The first four parts covered prefill and decode pools, the gateway, router and autoscaler, hardware failures and rollouts. They mostly assumed one kind of traffic, a user chatting. A coding agent, a reasoning model and an overnight batch job send very different requests, and some of the machinery from those posts matters less for them, or more.

This part describes a workload with a few numbers, goes through the common types, and then goes back through parts 1 to 4 to see which workloads each piece is for. The last section measures, on one H100, how much a few long prompts slow down everyone else’s answers, and which latency numbers show it.

The numbers that describe a workload Link to heading

Six numbers cover most of it:

  • Input and output length. Long inputs make prefill the expensive part; long outputs make it decode. Many benchmarks name a workload by the pair: InferenceMAX launched with 1k/1k for chat, 1k/8k for reasoning and 8k/1k for summarization. The ratio decides whether splitting prefill from decode (part 1) pays off, and how many GPUs each pool gets.
  • Latency target. A user waiting on screen, or a job that has until tomorrow. The gateway in part 2 uses it to decide whose request waits.
  • Prefix reuse. How much of each prompt some worker already has in its KV cache. It sets how much part 2’s router can save.
  • Context length. The KV cache grows with it, which limits how many requests fit in a decode batch, and how long it takes to prefill again after a failure (part 3).
  • Arrival pattern. Steady, bursty, or peaks at the same hours every day. In Azure’s 2024 traces, the busiest hour of the day had about 9 times the requests of the quietest for code, and about 1.8 times for conversation. The autoscaler in part 2 follows it, and a new replica takes minutes to arrive.
  • Request duration. How long a request holds a slot from start to finish. It sets the drain timeout in part 4 and how much a failure throws away.

Measuring latency Link to heading

Serving benchmarks report a request’s latency in two parts. TTFT (time to first token) is mostly queueing plus prefill. After that, each token takes about the same time, so:

end-to-end latency ≈ TTFT + TPOT × (output tokens - 1)

TPOT (time per output token) is one number per request, the average gap between its tokens. ITL (inter-token latency) records every gap as its own sample. The averages come out about the same, but the tails don’t. If a long prefill stalls decode for 500 ms once during a 1,000-token answer, that request’s TPOT goes up by half a millisecond, while its ITL gets one 500 ms sample in the tail. Which percentile shows it depends on how often stalls happen: in the measurement below they were just under 1% of gaps, so p99 barely moved and p99.9 caught them. The user sees the stream freeze. That’s the stall disaggregation part 1 measured.

These are the definitions in vLLM’s benchmark: TPOT is a request’s latency minus its TTFT, divided by its output tokens minus one, and ITL is every gap between streamed chunks, from all requests, in one list. A chunk can hold more than one token, which happens with speculative decoding, so with it on ITL counts gaps between chunks, and TPOT is the steadier number.

Batch size trades one user’s TPOT against the GPU’s total output. Each decode step reads all the model’s weights once, however many requests are in the batch, so adding requests makes each step a little slower and produces more tokens per step. Part 3 has an example by accident: when replica B went from 8 answers to 16 after the failure, its median gap between tokens went from 6.5 ms to 7.4 ms. Each user got tokens about 14% slower, and B as a whole went from about 1,230 tokens per second to about 2,160. Long contexts limit how far this goes, because every request in the batch needs its KV cache in GPU memory.

Common workloads Link to heading

A few providers have published traces of real requests, and they give the input and output lengths. I computed these from the raw files: Azure’s 2024 traces of its LLM services (one week each), Kimi’s Mooncake traces (one hour), and BurstGPT, from ChatGPT and GPT-4 traffic through Azure. Token counts are per request; the last column is total input tokens over total output tokens.

trace                             requests   input p50 / p90    output p50 / p90   input:output
Azure, conversation               27.3M      928 / 3,830        41 / 342           15:1
Azure, code                       16.8M      1,930 / 6,251      8 / 43             111:1
Kimi, conversation                12,031     6,909 / 27,367     350 / 597          35:1
Kimi, tool and agent              23,608     6,346 / 16,804     30 / 507           47:1
BurstGPT, ChatGPT conversation    310,015    468 / 1,772        168 / 484          3.5:1
BurstGPT, GPT-4 conversation      146,898    687 / 2,736        319 / 795          3.0:1

Even chat sends more tokens in than out, by 3 to 35 times depending on the service. Azure doesn’t say what its code service is, but a median answer of 8 tokens looks like completion more than chat.

Workload Input / output Latency target Prefix reuse Notes
Chat hundreds to a few thousand, grows each turn / a few hundred at most TTFT, then ITL about reading speed high (earlier turns) what most serving setups are tuned for
Code autocomplete about 2k (the open file) / tens very tight TTFT high (same file) about half the requests are cancelled
Agents over 100k by the end of a task / a few hundred per call total time across many calls very high idle gaps while tools run and users think
Reasoning short / thousands to tens of thousands ITL more than TTFT low to medium mostly decode, long-running requests
RAG and summarization long / short TTFT low unless the same documents come back mostly prefill
Offline batch any none, hours are fine varies can be paused or preempted, fills quiet hours

Embeddings and rerankers are prefill only: the model reads the input, returns a vector or a score, and keeps no KV cache. Multimodal models add a step before prefill, where an encoder turns each image or video into tokens.

Autocomplete throws a lot of work away. A completion is out of date as soon as the user types another character, so the client cancels it. GitHub’s David Cheney said that for Copilot “cancellation occurs every other request on average”, and Meta found that “47% of all suggestions generated by the model were never displayed” before it added streaming with cancellation.

Agents need more than a table row. Each call resends the whole conversation so far plus whatever the last tool returned. In traces of GitHub’s Copilot coding agent, 13 million sessions over a week in June 2026, the median session made 15 model calls, the median call produced 247 tokens, and the median input-to-output ratio was over 275:1. In TraceLab, which logged about 4,300 Claude Code and Codex sessions from 43 developers, “a median Claude step reads back 126k prefix tokens but appends only 857”. Nearly all of that is cached: Copilot’s agent hit the cache about 98% of the time at the median, and TraceLab measured 95.7%. TraceLab found most misses came between turns, when the user took longer to read and reply than the cache survived. Between calls, the worker holding an agent’s cache has to decide whether to keep it, move it to CPU memory, or drop it.

Reasoning models go the other way. The prompt is often short and the answer is thousands of thinking tokens: QwQ-32B averaged 2,408 output tokens on MATH500 and 9,481 on AIME24 in one measurement, and OpenAI recommends “reserving at least 25,000 tokens for reasoning and outputs”. Most of the time is decode, and each request holds its slot far longer than a chat reply.

Back through the series Link to heading

Splitting prefill and decode Link to heading

Part 1 gave two reasons to split, and both depend on the input and output lengths. The stall happens when a long prompt arrives at a GPU that is decoding, which describes RAG and summarization traffic, and agent calls that miss the cache. Eric Zhang’s model puts the speedup at 1 divided by the share of time a shared engine spends decoding. That’s large for long prompts with short answers, and close to 1 for reasoning models, which already spend most of their time decoding. Offline batch has no ITL target, so a stall costs it nothing it measures. As far as I can tell, splitting helps batch only if it gets more tokens out of each GPU-hour, which depends on keeping both pools busy.

Prefix-aware routing Link to heading

Part 2’s router saves work only when a new request starts with a prefix some worker already holds. In llm-d’s benchmark, 150 customers sharing 6,000-token contexts, routing on precise cache events cut p90 time to first token from 93 s (random) to 0.54 s. Chat and agents look like that benchmark, since each call resends the earlier turns, and an agent can make many calls on the same growing context within one task. Kimi’s tool and agent trace marks each 512-token block of a prompt with a hash, and 55% of its blocks had already appeared earlier in the same hour. That’s the best a router could do with unlimited cache; a real fleet evicts some of them first. RAG over different documents per request and batch jobs with unrelated prompts give the router little to find, and scoring on load alone should do about as well.

Gateway priority and autoscaling Link to heading

Chat and batch sharing the same GPUs is the reason the gateway has priorities. In an earlier design of the Gateway API Inference Extension, a Sheddable request only went to servers under 80% KV cache use with fewer than 5 requests waiting, so batch work ran in the room interactive traffic left. DeepSeek does the same on a daily cycle, serving inference on all its nodes during the day and handing nodes to research at night.

The mix also decides which autoscaling signal works. A count of waiting requests treats “Hi” and a 30,000-token document the same, so it under-counts the work when RAG prompts get longer. Dynamo’s planner sizes prefill from input tokens per second and decode from output tokens per second, so more reasoning traffic grows the decode pool and more RAG traffic grows the prefill pool.

Restart or continue after a failure Link to heading

In part 3, replica A died after 128 tokens of a 512-token answer. Restart and continue both had to prefill the prompt again on B, and continue then saved generating those 128 tokens a second time: B finished 0.7 to 0.8 s sooner. A reasoning answer that fails at token 6,000 raises both costs. Restarting means decoding 6,000 tokens again, about 40 seconds at the 6.5 ms per token B managed with 8 users. Continuing means prefilling 6,000 more tokens, which part 3’s numbers put well under a second (8,692 tokens took 0.23 s), over a longer context, where part 3 also saw the continued answer drift from the original at 28k tokens. For a chat reply of a few hundred tokens, the choice barely matters.

Drain timeouts Link to heading

llm-d says to set --shutdown-timeout to about the p99 request duration (part 4). Part 4’s five 2,048-token answers drained in about 10 seconds. Each agent call is short, but cutting one off stalls the whole task until the agent’s client retries it. Reasoning traffic is where the timeout runs out. OpenAI says reasoning models “can take several minutes”, which is longer than the default grace periods in part 4 (30 seconds for Kubernetes, 60 for Dynamo’s pods). A fleet serving them needs part 3’s request migration or clients that retry, on top of a longer drain.

Measurement Link to heading

I ran vLLM 0.30 with Qwen2.5-7B-Instruct on one Modal H100, the same setup as parts 2 to 4, with prefix caching off. The load came from vLLM’s own vllm bench serve with random prompts, 16 requests at a time, every answer forced to its full length. Each measurement ran twice. The script and the raw results are in the lab repo.

First, three request shapes on their own, with chunked prefill off:

input / output    TTFT      mean TPOT   p99 ITL   max ITL
8,192 / 256       1.21 s    15.9 ms     195 ms    206-214 ms
1,024 / 1,024     0.23 s    6.8 ms      7.6-7.8   154-170
256 / 4,096       0.09 s    7.0 ms      8.1-8.2   54

When every request has a long prompt, nothing is hidden. Some request is always prefilling, so the mean gap between tokens more than doubles and p99 is a 195 ms stall.

Traffic is usually a mix, so the second test ran a decode-heavy stream (256-token prompts, 1,024-token answers, 16 at a time) alone, and then again while a second client sent one 8,192-token prompt per second, each with a 16-token answer. These are the decode stream’s numbers:

                          mean TPOT   p50 ITL   p90 ITL   p99 ITL    p99.9 ITL   max ITL
alone                     6.7-6.9 ms  6.7 ms    7.0-7.7   7.3-15     8-26        14-66
chunking off              8.3         6.8       7.3       12.1-12.4  190-191     192-194
8,192-token chunks        8.3         6.7       7.1       12.8       186         187-192
2,048-token chunks        8.1-8.2     6.7       7.1-7.2   53.6-53.7  55.9        57-58
512-token chunks          7.7         6.8       13.3-13.5 16.0       16.5-17.0   19

With chunking off, every 8k prompt froze all 16 streams for about 185 ms. That happened once a second, and mean TPOT rose 23%, from 6.7 ms to 8.3: about 1.6 seconds of stalls spread over each 1,024-token answer. The median didn’t move. A second has about 1,900 gaps across the 16 streams, and 16 of them were stalls, under 1%, so p99 only went from about 8 ms to 12. p99.9 and max show the 190 ms freeze. vLLM 0.30’s default budget of 8,192 tokens runs almost all of an 8k prompt in one step, and it looked the same as chunking off.

Smaller chunks did what they did in disaggregation part 1: one long stall became several short ones. With 2,048-token chunks, each 8k prompt took four steps of about 54 ms, enough of the gaps to set p99. With 512-token chunks it took sixteen steps of about 13 ms, and those moved p90, while the worst gap fell to 19 ms and mean TPOT was the lowest of the four at 7.7 ms. The chunk size decides which percentile carries the cost; none of the settings removed it, which is part 1’s case for giving prefill its own GPUs.

Caveats Link to heading

  • One model and one GPU type, with synthetic request shapes in place of real traces.
  • The traces are samples. Azure’s 2024 inputs stop just under 8,000 tokens, which looks like a context cap, so their long tail is cut off. Kimi’s cover one hour. None of them says which model served the requests.
  • The 8k prompts’ rate comes from the second client’s own progress log, about 1.03 completed per second. vLLM’s prompt-token counter suggested about 50 per run instead of about 35; the TPOT arithmetic above matches one per second, so I used that.
  • vLLM 0.30 warns that this model “does not officially support disabling chunked prefill”. The longer TTFT and 195 ms stalls with it off suggest the setting took effect.
  • Real fleets mix these workloads on the same GPUs. The types above are a way to describe traffic, and most traffic is some mix of them.