“How many GPUs does it take to serve this model to these users?” comes up in every capacity plan. A handful of formulas get you to an estimate. This post works one problem from start to finish and introduces each formula where it’s needed. Each one is checked against numbers measured in Weight precision part 1 and part 2: Qwen3-8B in BF16 on vLLM 0.30, one run on one H100 and one B200.
Serve Qwen3-8B in BF16. At peak, 20 requests arrive per second. Prompts are 2,000 tokens and answers are 400 tokens. Each user should get at least 50 tokens per second, so at most 20 ms per token, and the first token within a second. How many H100s, or how many B200s?
1. Requests in flight Link to heading
Traffic arrives as a rate, but memory and latency depend on how many requests are in the system at once. Little’s law connects them:
requests in flight = arrival rate × time each request spends in the system
time per request = 0.2 s first token + 400 × 0.020 s = 8.2 s
in flight = 20 per second × 8.2 s = 164 requests
2. Memory per replica Link to heading
A replica’s GPU memory holds the weights, the KV cache, and the engine’s own working memory:
weights = parameters × bytes per weight
KV per token = 2 (K and V) × layers × KV heads × numbers per head × bytes per number
The engine decides how much of that goes to the KV cache, at startup. vLLM 0.30 does it in four steps:
- It takes a share of GPU memory, set by
--gpu-memory-utilization. The default is 0.92, so 8% stays free for anything else on the GPU. - It loads the weights.
- It runs a test forward pass to measure its working memory: the activations of the largest batch it allows, plus the CUDA graphs it records.
- Whatever is left becomes the KV cache, divided into blocks of 16 tokens.
For Qwen3-8B on an 80 GB H100, from the startup log in Weight precision part 1:
GPU memory 79.6 GiB
vLLM's share, × 0.92 73.3 GiB
− weights (8.19 billion × 2 bytes) 15.3 GiB
− working memory 4.6 GiB
= KV cache 53.4 GiB
KV per token: 2 × 36 × 8 × 128 × 2 bytes = 147,456 bytes
53.4 GiB / 147,456 bytes = 388,560 tokens (24,285 blocks of 16)
The B200 works the same way: 179.1 GiB × 0.92, minus the same 15.3 GiB of weights and 4.3 GiB of working memory, leaves 145.2 GiB, or 1,056,992 tokens. So the KV room depends on the GPU, the weight format, the KV cache format and the memory share. It also depends on --max-model-len, which sets how large a batch the working memory has to cover. Change any of them and the number of requests that fit changes with it.
Each request in this problem holds its prompt and answer in the cache:
KV per request = (2,000 + 400) tokens × 147,456 bytes = 354 MB
H100: 57.3 GB of KV room / 354 MB = 161 requests
B200: 1,056,992 tokens of room / 2,400 = 440 requests
That’s how many requests fit in a replica: 161 on the H100 and 440 on the B200. Fitting isn’t the same as serving them fast enough. Each one added makes every user’s tokens a little slower, so the 20 ms target allows fewer, as the next section works out.
3. Latency per replica Link to heading
Each decode step reads every weight once, however many requests share the step, plus each request’s KV cache. With small batches, reading is the slow part:
time per step ≈ bytes read per step / memory bandwidth + fixed cost
A decode step reads 15.1 GB of Qwen3-8B’s 16.4 GB of weights. It skips most of the other 1.24 GB, the embedding table. That’s the model’s first layer. The layers after it do arithmetic on vectors, and a token ID is only a position in the vocabulary, so the table swaps each ID for a learned vector of 4,096 numbers. It has one row for each of the 151,936 tokens in the vocabulary. A step needs only the row for each request’s latest token. On the H100 at 3.35 TB/s, that predicts 4.5 ms with one user. The measured time was 6.6 ms. Comparing that BF16 run with an FP8 run, which reads 8.3 GB per step, separates the two parts:
BF16: 6.60 ms = 15.1 GB / bandwidth + fixed
FP8: 4.51 ms = 8.3 GB / bandwidth + fixed
subtract: 2.09 ms = 6.8 GB / bandwidth → bandwidth = 3.3 TB/s, fixed = 2.0 ms
The 2 ms is attention, sampling, launching kernels and the scheduler. Under load, the formula gets optimistic. With 16 requests at 8k context, the predicted 12.5 ms came out as 17.1 ms measured: reading a large, scattered KV cache ran at about 2.3 TB/s, not 3.35. The B200’s equivalent was 5.6 TB/s. So calibrate the bandwidth once, at a load like yours.
Why does adding requests cost so little? Each request adds 2 FLOPs per weight, and in BF16 each weight is 2 bytes: 1 FLOP per byte read. An H100 can do 989 TFLOPS / 3.35 TB/s = 295 FLOPs per byte, so FLOPs don’t become the limit until about 295 requests share a step:
On the B200, a step took 4.07 ms with 1 request and 4.26 ms with 16 short ones: 16 times the tokens for 5% more time (raw results). That’s the second of two runs. The first came out at 7.25 ms, which I took for a warmup effect but didn’t confirm. What does grow with each request is its KV cache. In this problem a request averages 2,200 tokens of context during decode, so it adds 0.32 GB per step:
time per step = (15.1 GB + 0.32 GB × requests) / effective bandwidth + fixed cost
H100: (15.1 + 0.32 × r) / 2.3 TB/s + 2.0 ms ≤ 20 ms → r ≤ 81
B200: (15.1 + 0.32 × r) / 5.6 TB/s + 1.9 ms ≤ 20 ms → r ≤ 266
That suggests 164 / 81 = 3 H100 replicas, or 1 B200. It’s too few.
4. Prefill takes time from decode Link to heading
Every new request needs its 2,000-token prompt prefilled, and while a GPU prefills, its decoding users wait. Inference in production part 5 measured that stall. Prefill is limited by FLOPs, not bytes. Each weight is read once and used for every token in the prompt, a multiply and an add each. A 2,000-token prompt does about 2,000 FLOPs per byte read, far past the 295 where the GPU runs out of FLOPs before memory.
FLOPs ≈ 2 × weights multiplied × prompt tokens + attention
attention ≈ 2 × layers × prompt tokens² × model width (with a causal mask)
The embedding table is a lookup, not a multiply, so Qwen3-8B multiplies 7.57 billion weights. Checked against the measured 8k-token prompts:
8,192 tokens: 2 × 7.57 billion × 8,192 + 2 × 36 × 8,192² × 4,096 = 144 TFLOP
H100 at 989 TFLOPS (BF16, dense): 146 ms at peak measured: 250 ms → 58% of peak
B200 at 2,250 TFLOPS: 64 ms at peak measured: 136 ms → 47% of peak
Real kernels reached about half of the datasheet’s peak. At those shares, this problem’s 2,000-token prompts take:
2 × 7.57 billion × 2,000 = 30.3 TFLOP → H100: 53 ms B200: 29 ms
With N replicas, each gets 20 / N requests per second and spends that many prefills’ worth of each second not decoding. Decode gets what’s left of the 20 ms budget:
H100, N = 3: 6.7 per second × 53 ms = 35% prefilling → 13.0 ms for decode → r ≤ 31 → 93 total too few
H100, N = 4: 5.0 per second × 53 ms = 26% prefilling → 14.7 ms for decode → r ≤ 43 → 172 total enough
B200, N = 1: 20 per second × 29 ms = 57% prefilling → 8.5 ms for decode → r ≤ 68 → 68 total too few
B200, N = 2: 10 per second × 29 ms = 29% prefilling → 14.3 ms for decode → r ≤ 166 → 332 total enough
Counting prefill added a replica on each GPU type. Memory isn’t the limit: 43 and 166 requests are well under the 161 and 440 that fit.
5. GPUs Link to heading
Add one spare for a failure and one for a rolling update, as in Inference engines part 2:
H100: 4 replicas + 2 spares = 6 GPUs
B200: 2 replicas + 2 spares = 4 GPUs
The B200 fleet is cheaper whenever a B200 costs less than 6 / 4 = 1.5 times an H100 per hour. Plug in your own prices.
The first-token target likely holds: a prefill takes 53 ms on the H100, and each replica is prefilling about a quarter of the time, so most new requests wait little. Confirming the p99 needs a load test.
Where the real numbers may differ Link to heading
- The calibration is one run. Effective bandwidth, fixed cost and share of peak came from one H100 and one B200, at 16 requests with 8k contexts. Re-check the answer with a load test at the planned size.
- Real prompts vary. A few long prompts take a bigger share of prefill time than the average suggests.
- Averages hide bursts. Little’s law gives the average in flight. Traffic that spikes above 20 per second needs headroom, or the spares get used for load instead of failures.
- Prefill and decode share the GPU more cleverly than this. Chunked prefill interleaves them, so the “share of time” model is rough.
A 70B variant Link to heading
For a 70B model, the same five steps start from much larger weights:
weights in BF16: 70 billion × 2 bytes = 140 GB → at least 2 H100s just to hold them
→ 4 to leave room for the cache
decode floor, 4 H100s: 140 GB / (4 × 3.35 TB/s) = 10.4 ms per step, before any cache or fixed cost
That floor already uses half of a 20 ms target, so each replica fits far fewer requests, and each replica is 4 GPUs. In FP8, the weights take 70 GB on 2 H100s, with the same 10.4 ms floor on half the GPUs. The KV per token comes from that model’s config, and the rest follows the same steps.