Skip to content
GenAI Learn/Diffusion Models & Multimodal Generation
Browsing as a guest. Sign in to save your progress and earn XP as you complete chapters.

Audio, Speech & Where Multimodal Is Headed

8 min read

You'll learn to

  • -Explain why speech models process a spectrogram, not a raw waveform
  • -Understand the Discrete Fourier Transform as the mathematical basis of every spectrogram
  • -Describe, at a conceptual level, how video generation and native multimodal models extend the ideas from this module

Text and images are not the only modalities a generative system deals with. This closing chapter covers the one piece needed to make sense of speech models (Whisper-style transcription, modern text-to-speech), and previews where multimodal generation is headed next: video, and models that handle every modality natively instead of stitching specialists together.

How Machines "Hear": From Waveform to Spectrogram

A raw audio waveform is just a long sequence of amplitude numbers - not a useful representation for a neural network to reason over directly. Every speech model instead works on a spectrogram: a picture of which frequencies are present in the sound, over time. The math underneath every spectrogram is the Discrete Fourier Transform (DFT): decompose a signal into how much of each frequency it contains.

Discrete Fourier Transform (magnitude)
real[k] = Σₙ signal[n]·cos(2πkn/N), imag[k] = -Σₙ signal[n]·sin(2πkn/N), |X[k]| = √(real[k]² + imag[k]²)
Real production systems use the FFT, a fast algorithm for computing the exact same result - the definition itself needs nothing but sums of sines and cosines.
The DFT magnitude spectrum, computed directly from its definition

Speech Models Built on Top: ASR and TTS

Whisper-style automatic speech recognition (ASR) feeds a spectrogram into an encoder-decoder Transformer - structurally, the same architecture family from the Transformers module, just with a spectrogram instead of token embeddings as the encoder's input. Text-to-speech (TTS) runs conceptually in reverse: a model generates a spectrogram conditioned on text, and a separate model (a vocoder) turns that spectrogram back into an audible waveform - often using a diffusion process much like the one from this module's first chapter, just generating audio frames instead of image pixels.

Video Generation & Where Multimodal Is Headed

Video generation (Sora and similar systems) extends this module's diffusion process one dimension further: instead of denoising a 2D grid of pixels, these models denoise "spacetime patches" - small chunks of video spanning both space and time - using a Transformer-based denoiser instead of the U-Net architecture classic image diffusion models use, but the same forward-noise/reverse-denoise process from this module's first chapter underneath.

Two Ways to Build a Multimodal System
Separate models for each modality (an ASR model, an LLM, a TTS model, a diffusion model), glued together - easier to build and swap pieces, but errors compound across the pipeline and latency stacks up
Pipeline of specialists
One model trained to understand and generate across modalities directly (the GPT-4o/Gemini direction) - harder to train, but avoids the information loss and latency of converting between modalities at every pipeline boundary
Native multimodal model

AI Lab's Level 49 builds the DFT magnitude spectrum in numpy, vectorized across every frequency bin at once with a single matrix multiply instead of the nested loop above.

Interview Signal is part of Pro

See a real weak answer next to a real strong one for this exact topic.

Quiz is part of Pro

Test what you just read with a short quiz, and bank the XP.

Ready to Build This?

Implement it from scratch: Level 72 (Mel-Style Spectrograms via a from-scratch DFT) in AI Lab.