Inference Capacity Planning: Quantization & Mixed-SLA Pools
You'll learn to
- -Distinguish KV-cache quantization from model quantization, and know when the trade-off is worth it
- -Reason about why one shared GPU pool fails workloads with different latency SLAs
- -Compute the real capacity gain from quantizing the KV-cache, instead of just assuming "roughly double"
The first chapter in this module fixed serving for one workload: continuous batching and PagedAttention keep a single GPU pool busy and efficient. Real companies usually run two workloads through the same model at once, though, and they rarely have the same SLA. A live chat feature needs a response in well under a second. An overnight batch job rewriting ten thousand documents does not care if it takes two hours or four, as long as it is cheap. Pointing both at one undifferentiated pool, sized for whichever workload someone thought about first, is the step most teams skip, and it shows up as either wasted GPU-hours or a latency incident.
KV-Cache Quantization: The Lever Batching Alone Does Not Give You
Continuous batching schedules requests efficiently, but it does not change how much GPU memory each one consumes. The KV-cache for a request is normally stored at fp16 precision. Quantizing it to fp8 or int8 roughly halves the memory each sequence needs, which means roughly double the number of concurrent sequences fit in the same GPU memory. This is a different knob from quantizing the model's weights (which trades accuracy for a smaller model footprint) - here, the model runs at full precision, only the per-request cache shrinks.
That 2x figure holds at any context length, since both precisions scale the same way as the prompt gets longer. Drag the slider below to see it recompute live.
Concurrent Capacity on One A100-80GB, Same Context Length
Longer context means a bigger KV-cache per sequence, so fewer sequences fit at once. Quantizing that cache to fp8 roughly halves its size, so roughly double the sequences fit in the same 80GB. Drag the context length and watch both capacities recompute live.
fp16 KV-cache (no quantization)
65 sequences
fp8 KV-cache (quantized)
130 sequences
At 2,048-token context, quantizing the KV-cache alone buys 2.0x the concurrent capacity on the same GPU - the right trade for a deferrable throughput pool, not for a latency-critical one.
That capacity gain is not free, and it is not the right call everywhere. Lower-precision KV-cache means slightly less numerical precision in attention, which can show up as subtly different outputs. For an overnight summarization job where nobody is comparing token-for-token against an fp16 run, that is a non-issue. For a latency- and quality-critical interactive feature, it is a real trade a team should make deliberately, not inherit by accident because one pool's config got copied everywhere.
Why One Pool Fails Two SLAs
Size the shared pool for the chat workload's peak, and it sits mostly idle overnight while the batch job trickles through far slower than the hardware allows. Size it for the batch job's throughput instead, large batch size, quantized KV-cache, and an unlucky long-document request queued ahead of a chat request can blow straight through the chat SLA the moment both workloads overlap. Neither config is wrong in isolation. The mistake is expecting one pool, one config, to be simultaneously latency-optimized and throughput-optimized, which are opposite ends of the same batch-size and precision knobs.
Routing by Complexity, Not Round Robin
- -Classify incoming requests by a cheap signal available before generation starts: prompt token count, a declared urgency flag, or which endpoint the request hit.
- -Route short, urgent requests to a latency pool: smaller batch size, fp16 KV-cache, headroom deliberately left unused so a burst never queues.
- -Route long, deferrable requests to a throughput pool: large batch size, fp8 or int8 KV-cache, tuned to pack the GPU as tightly as the SLA allows.
- -Track cost per pool separately. A blended, single cost-per-request number hides which pool is actually efficient and which one is quietly over budget.
This is the same InferenceRouter component from the gateways chapter, just classifying on request shape instead of provider health or price. The routing logic does not change; what it is optimizing for does.
The shape the Capacity Planner level asks you to wire up: a router classifying traffic before it ever reaches a pool, with cost and latency tracked per request after.
Interview Signal is part of Pro
See a real weak answer next to a real strong one for this exact topic.
Quiz is part of Pro
Test what you just read with a short quiz, and bank the XP.
Build the Serving Engine, Gateway, Cache Layer, and Capacity Planner levels in the GenAI Lab.
Discussion0
Join the Discussion
Sign in to leave comments, reply to others, or like insights.
No comments yet. Be the first to start the thread!