Skip to content
GenAI Learn/Transformers Under the Hood
Browsing as a guest. Sign in to save your progress and earn XP as you complete chapters.

Scaling Laws, Quantization & Mixture-of-Experts

9 min read

You'll learn to

  • -Compute the compute-optimal parameter/token split for a training budget using the Chinchilla scaling law
  • -Implement symmetric int8 quantization and explain the memory/accuracy tradeoff it makes
  • -Understand Mixture-of-Experts top-k routing as the mechanism behind sparse, larger-than-they-look models like DeepSeek and Mixtral

Everything so far in this module explains HOW a Transformer block computes its output. This chapter is about three decisions that shape WHICH model you are even running: how big it should be for a given training budget, how cheaply it can be served once trained, and how it can be made bigger without a proportional increase in per-token compute.

Scaling Laws: Why a Model Is the Size It Is

Given a fixed training compute budget, should you train a bigger model on less data, or a smaller model on more data? DeepMind's Chinchilla paper answered this empirically: compute-optimal training grows model size (N parameters) and dataset size (D tokens) together, in a roughly fixed ratio of about 20 training tokens per parameter. GPT-3-era models were trained well short of that ratio - undertrained relative to their size - which is a big part of why smaller, longer-trained models like Llama outperform much larger, older ones.

Compute-Optimal Allocation
C = 6ND ⟹ N = √(C / (6·ratio)), D = ratio·N
Training compute follows C ≈ 6 FLOPs per parameter per token. Substituting the Chinchilla ratio D = ratio·N and solving for N gives the compute-optimal split.
Solving for the compute-optimal model size and dataset size

Quantization: Serving a Model for a Quarter of the Memory

A model's weights are normally stored as 16- or 32-bit floats. Quantization stores them as 8-bit integers instead - a 2-4x memory reduction, and often a real speedup since integer arithmetic is cheaper - at the cost of a small, controlled rounding error. This is the core idea behind GPTQ and AWQ, the two most common ways production LLMs get compressed for cheaper serving.

Symmetric int8 Quantization
scale = max(|w|) / 127, q = round(w / scale), ŵ = q · scale
The largest-magnitude weight maps to the edge of the int8 range; every other weight is stored as a small integer plus that one shared scale factor.
Quantizing a small batch of weights to int8, and back

Mixture-of-Experts: A Bigger Model at the Same Per-Token Cost

A dense model activates every parameter on every token - doubling model size doubles compute per token. Mixture-of-Experts (MoE) breaks that link: a layer holds many "expert" sub-networks, but a lightweight router activates only the top few for each token. DeepSeek-V3 has over 600 billion total parameters but activates around 37 billion per token - a huge model at small-model inference cost. The router is the whole trick.

  • -A router produces one logit per expert per token - like a classifier choosing which expert(s) should handle this token, not which class it belongs to.
  • -Top-k gating keeps only the k highest logits per token and turns those into a softmax distribution renormalized among just the k survivors - every other expert gets a gate weight of exactly 0.
  • -A gate weight of 0 means that expert is never even called for that token - compute scales with k (active experts), not with the total number of experts in the layer.
Top-k routing: only 2 of 4 experts get called for this token

AI Lab's Levels 40-42 implement all three of these in numpy: solving the Chinchilla equations directly, real int8 quantize/dequantize round-tripping, and vectorized top-k gating across a whole batch of tokens at once.

Interview Signal is part of Pro

See a real weak answer next to a real strong one for this exact topic.

Quiz is part of Pro

Test what you just read with a short quiz, and bank the XP.

Ready to Build This?

Implement all three: Level 63 (Chinchilla Scaling Law), Level 64 (8-bit Quantization), and Level 65 (Mixture-of-Experts Routing) in AI Lab.