Ashok Popat
AshokAudio-native representations for voice interaction; OCR, handwriting, and graphics understanding. Background in data compression and statistical NLP.
Learning path · for Ananya Modugula · September 2026
From training language models to working on voice interaction, audio representations, streaming generative models, and symbol recognition, which is the work of Ashok Popat and RJ Skerry-Ryan at Google DeepMind.
How do you turn a continuous signal, such as speech or a pen stroke, into something a language model can reason over, and turn it back into a signal, in real time?
Ashok works on the representation side. RJ works on generation and real-time systems. You've already trained a transformer from scratch; what's missing is the signal side.
Audio-native representations for voice interaction; OCR, handwriting, and graphics understanding. Background in data compression and statistical NLP.
Co-inventor of Tacotron. Full-duplex dialogue for Gemini Live and Astra, streaming generative models, and the SequenceLayers library.
About 12 to 15 months at 8 to 10 hours a week
torchaudio.