LLM Inference & Serving Capacity

GenAI Systems
22 / 22

Serving an LLM in production is a capacity-planning problem, not a model-choice problem: batching, KV-cache memory, and GPU fleet sizing decide your real throughput and cost per token.

The Analogy (Read This First)

A restaurant doesn't need a faster chef to serve more tables - it needs a bigger pass and an efficient queue. LLM serving is the same: PagedAttention manages the "pass" (GPU memory for in-flight requests), continuous batching manages the queue, and the GPU/quantization choice decides how many "tables" you can seat at once.

Deep Dive Analysis

Every LLM request holds a KV-cache in GPU memory for every token it's generating - the longer the prompt and the more concurrent requests, the more memory gets consumed. Naive serving allocates a fixed, worst-case memory block per request, wasting most of it most of the time. vLLM's PagedAttention (2023) fixed this by paging KV-cache like an OS pages virtual memory - allocating it in small blocks on demand - which alone unlocked 2-4x more concurrent requests on the same GPU.

Continuous batching (vs static batching) is the second lever: instead of waiting for an entire batch to finish before starting the next one, a finished request's slot is immediately handed to a new request already waiting. On mixed-length traffic (some 50-token replies, some 2,000-token documents), this alone can 3-5x throughput because short requests stop waiting behind long ones.

Quantizing the KV-cache itself (fp16 → fp8 or int8) roughly halves the memory each sequence needs, which lets you raise the batch size further - a real throughput lever for cost-sensitive, latency-tolerant traffic (batch summarization, overnight jobs), though it is the wrong trade for latency-critical interactive traffic where precision and every millisecond matter.

None of this is one-size-fits-all: a chat product with a tight p95 SLA and a batch-processing job with none should almost never share one undifferentiated GPU pool. Routing by request complexity (prompt length, urgency, SLA) into separate pools - one tuned for latency (small batch, fp16, headroom to spare), one tuned for throughput (large batch, fp8/int8 KV-cache, packed tight) - is how teams serving both workloads hit both their latency and cost targets at once.

Key Bullets
  • PagedAttention pages KV-cache like OS virtual memory, eliminating the waste of fixed-size pre-allocation per request.
  • Continuous batching lets a finished request's slot go straight to a waiting one, instead of waiting for the whole batch to drain.
  • KV-cache quantization (fp8/int8) trades a little numerical precision for roughly double the concurrent sequences per GPU - the right trade for throughput-first, not latency-first, traffic.
  • Capacity planning means matching GPU pool configuration to the SLA of the workload, not running one generic pool for everything.
Trade-offs
✅ Continuous batching + PagedAttention: 2-5x more throughput on the same hardware, no model changes required✅ KV-cache quantization: roughly doubles concurrent capacity for cost-sensitive, latency-tolerant workloads❌ Quantization is the wrong call for latency-critical interactive traffic - don't apply it uniformly❌ Splitting pools by SLA adds real operational complexity: more configs, more routing logic, more ways to misroute a request
Real-World Examples
vLLM: the reference open-source implementation of PagedAttention and continuous batchingNVIDIA TensorRT-LLM and Hugging Face TGI: production serving frameworks built on the same batching/caching principlesScaleDojo GenAI Lab, "The Capacity Planner": design a mixed-SLA serving fleet around exactly this trade-off

Curated Curation & Deep Insights

ScaleDojo Certified
Video Tutorial Pending

Our system architects are vetting high-quality, authorized video guides for LLM Inference & Serving Capacity with zero third-party platform links.

Verification in Progress
Reading Material Pending

We are preparing premium, zero-competitor deep-dives for LLM Inference & Serving Capacity. Authorized reading references will appear automatically.

Verification in Progress