Learning path · for Ananya Modugula · September 2026

Signal to symbol

From training language models to working on voice interaction, audio representations, streaming generative models, and symbol recognition, which is the work of Ashok Popat and RJ Skerry-Ryan at Google DeepMind.

The one question

How do you turn a continuous signal, such as speech or a pen stroke, into something a language model can reason over, and turn it back into a signal, in real time?

Ashok works on the representation side. RJ works on generation and real-time systems. You've already trained a transformer from scratch; what's missing is the signal side.

Where you stand

What carries over

  • Smol-Llama: speech LLMs are the same machine with audio tokens.
  • Statistics: latent-variable models and rigorous listening tests.
  • RL and full-stack skills: dialogue policy and real-time demos.
  • Already at Google: internal paths outside candidates don't have.

What's missing

  • Signal processing and speech science
  • VAEs, GANs, flows, diffusion, and flow matching
  • Audio tokenizers, streaming design, and JAX
  • A research record: reproductions and write-ups

The two researchers

Ashok Popat

Ashok

Audio-native representations for voice interaction; OCR, handwriting, and graphics understanding. Background in data compression and statistical NLP.

RJ Skerry-Ryan

RJ

Co-inventor of Tacotron. Full-duplex dialogue for Gemini Live and Astra, streaming generative models, and the SequenceLayers library.

Seven phases

About 12 to 15 months at 8 to 10 hours a week

1

Signals and speech fundamentals

Both~6 weeks
Learn
STFT and inverse STFT, mel spectrograms, phase, resampling, pitch, and prosody.
Read
Think DSP and the speech chapters of Jurafsky and Martin.
Build
A differentiable STFT, mel, and Griffin-Lim toolkit in PyTorch that matches torchaudio.
2

Neural speech synthesis and recognition

RJ~8 weeks
Learn
Attention alignment, prosody as a latent variable, vocoders, CTC, and RNN-T.
Read
Tacotron, Tacotron 2, prosody transfer, and HiFi-GAN.
Build
Tacotron 2 on LJSpeech with a style encoder, plus a page of audio samples.
3

Generative modeling toolbox

RJ~8 weeks
Learn
VAEs, GANs, normalizing flows, diffusion, flow matching, and few-step sampling.
Read
Stanford CS236, Flow matching, and Voicebox.
Build
One vocoder task solved four ways; compare quality, sampling steps, and latency.
4

Audio representations and tokenizers

AshokRJ~10 weeks
Learn
Self-supervised speech encoders, neural codecs, residual vector quantization, and rate-distortion.
Read
HuBERT, SoundStream, AudioLM, and Spectron.
Build
A small codec at three bitrates, then probe what each level encodes. The most research-shaped project here.
5

Full-duplex spoken dialogue

Both~10 weeks
Learn
Cascaded versus native speech models, turn-taking, interruptions, and latency budgets.
Read
dGSLM, AudioPaLM, and Moshi.
Build
A streaming voice agent with a measured latency breakdown, compared against Moshi.
6

Streaming models and JAX

RJ~8 weeks, alongside 4 and 5
Learn
Causal layers, state, lookahead, KV caches, and JAX on TPUs.
Read
The SequenceLayers report and How to Scale Your Model.
Build
Port your codec to SequenceLayers, prove streaming equals offline output, and open a pull request upstream.
7

Symbols in images

Ashok~8 weeks, any time after 2
Learn
OCR with CTC, OCR-free document models, chart understanding, and online handwriting.
Read
Pix2Struct, InkSight, and MathWriting.
Build
Rebuild TheCalc as a streaming handwritten-math-to-LaTeX recognizer.

Getting into the room