Audio, Speech & Where Multimodal Is Headed
You'll learn to
- -Explain why speech models process a spectrogram, not a raw waveform
- -Understand the Discrete Fourier Transform as the mathematical basis of every spectrogram
- -Describe, at a conceptual level, how video generation and native multimodal models extend the ideas from this module
Text and images are not the only modalities a generative system deals with. This closing chapter covers the one piece needed to make sense of speech models (Whisper-style transcription, modern text-to-speech), and previews where multimodal generation is headed next: video, and models that handle every modality natively instead of stitching specialists together.
How Machines "Hear": From Waveform to Spectrogram
A raw audio waveform is just a long sequence of amplitude numbers - not a useful representation for a neural network to reason over directly. Every speech model instead works on a spectrogram: a picture of which frequencies are present in the sound, over time. The math underneath every spectrogram is the Discrete Fourier Transform (DFT): decompose a signal into how much of each frequency it contains.
Speech Models Built on Top: ASR and TTS
Whisper-style automatic speech recognition (ASR) feeds a spectrogram into an encoder-decoder Transformer - structurally, the same architecture family from the Transformers module, just with a spectrogram instead of token embeddings as the encoder's input. Text-to-speech (TTS) runs conceptually in reverse: a model generates a spectrogram conditioned on text, and a separate model (a vocoder) turns that spectrogram back into an audible waveform - often using a diffusion process much like the one from this module's first chapter, just generating audio frames instead of image pixels.
Video Generation & Where Multimodal Is Headed
Video generation (Sora and similar systems) extends this module's diffusion process one dimension further: instead of denoising a 2D grid of pixels, these models denoise "spacetime patches" - small chunks of video spanning both space and time - using a Transformer-based denoiser instead of the U-Net architecture classic image diffusion models use, but the same forward-noise/reverse-denoise process from this module's first chapter underneath.
AI Lab's Level 49 builds the DFT magnitude spectrum in numpy, vectorized across every frequency bin at once with a single matrix multiply instead of the nested loop above.
Interview Signal is part of Pro
See a real weak answer next to a real strong one for this exact topic.
Quiz is part of Pro
Test what you just read with a short quiz, and bank the XP.
Implement it from scratch: Level 72 (Mel-Style Spectrograms via a from-scratch DFT) in AI Lab.