Scaling Laws, Quantization & Mixture-of-Experts
You'll learn to
- -Compute the compute-optimal parameter/token split for a training budget using the Chinchilla scaling law
- -Implement symmetric int8 quantization and explain the memory/accuracy tradeoff it makes
- -Understand Mixture-of-Experts top-k routing as the mechanism behind sparse, larger-than-they-look models like DeepSeek and Mixtral
Everything so far in this module explains HOW a Transformer block computes its output. This chapter is about three decisions that shape WHICH model you are even running: how big it should be for a given training budget, how cheaply it can be served once trained, and how it can be made bigger without a proportional increase in per-token compute.
Scaling Laws: Why a Model Is the Size It Is
Given a fixed training compute budget, should you train a bigger model on less data, or a smaller model on more data? DeepMind's Chinchilla paper answered this empirically: compute-optimal training grows model size (N parameters) and dataset size (D tokens) together, in a roughly fixed ratio of about 20 training tokens per parameter. GPT-3-era models were trained well short of that ratio - undertrained relative to their size - which is a big part of why smaller, longer-trained models like Llama outperform much larger, older ones.
Quantization: Serving a Model for a Quarter of the Memory
A model's weights are normally stored as 16- or 32-bit floats. Quantization stores them as 8-bit integers instead - a 2-4x memory reduction, and often a real speedup since integer arithmetic is cheaper - at the cost of a small, controlled rounding error. This is the core idea behind GPTQ and AWQ, the two most common ways production LLMs get compressed for cheaper serving.
Mixture-of-Experts: A Bigger Model at the Same Per-Token Cost
A dense model activates every parameter on every token - doubling model size doubles compute per token. Mixture-of-Experts (MoE) breaks that link: a layer holds many "expert" sub-networks, but a lightweight router activates only the top few for each token. DeepSeek-V3 has over 600 billion total parameters but activates around 37 billion per token - a huge model at small-model inference cost. The router is the whole trick.
- -A router produces one logit per expert per token - like a classifier choosing which expert(s) should handle this token, not which class it belongs to.
- -Top-k gating keeps only the k highest logits per token and turns those into a softmax distribution renormalized among just the k survivors - every other expert gets a gate weight of exactly 0.
- -A gate weight of 0 means that expert is never even called for that token - compute scales with k (active experts), not with the total number of experts in the layer.
AI Lab's Levels 40-42 implement all three of these in numpy: solving the Chinchilla equations directly, real int8 quantize/dequantize round-tripping, and vectorized top-k gating across a whole batch of tokens at once.
Interview Signal is part of Pro
See a real weak answer next to a real strong one for this exact topic.
Quiz is part of Pro
Test what you just read with a short quiz, and bank the XP.
Implement all three: Level 63 (Chinchilla Scaling Law), Level 64 (8-bit Quantization), and Level 65 (Mixture-of-Experts Routing) in AI Lab.