AI / ML / GenAI Engineer

Most of what I build listens for a while before it decides anything.

At Emitrr I shaved a third off the latency of a live voice pipeline — the kind of number that only matters because someone was mid-sentence when it did. Outside of work, I build small experiments that poke at the edges of these systems: a narrator that never quite explains itself, a classifier that has to explain everything, a model asked to find a few pixels of bleeding in fifty thousand frames.

Mumbai, India Open to full-time ML / AI engineering roles GitHub ↗ LinkedIn ↗

A rough shape of what I do

Production speech and LLM systems by day. Interpretable ML, agentic tooling, and the occasional game designed to make a language model slip, by whatever hours are left.

Emitrr — Speech Pipeline & LLM Systems Apr 2025 – Jan 2026

30% off transcription latency, a rebuilt TTS path, and an agent that learns what to do when a call falls apart.

Non-Invasive Blood Glucose Estimation Published research

A model that reads blood glucose out of a light-absorption waveform — one version accurate, one version an equation a person can actually read.

See all work →

Signal lens

Switch the lens to see how the same portfolio changes depending on the problem it's solving.

Speech systems

Latency focus 350 ms
Accuracy signal 96%
Core question Endpoint tuning and call-state recovery
System flow

Speech pipeline

Input Live audio
Model STT + endpoint logic
Decision Turn boundary + latency correction
Output Clean transcript / action