Part 1 followed one request through prefill and decode pools on B200 and GB200 NVL72. This part covers what sits in front of those pools: the gateway that accepts the request, the router that picks the GPUs for it, and the autoscaler that decides how many GPUs there are. The last section measures why a new replica takes minutes to become useful, with vLLM cold starts I timed on H100s. The rest comes from project docs and published benchmarks, linked where they’re used.
The gateway Link to heading
A request first reaches a gateway, the same as for any HTTP API: TLS, authentication, and a route picked by the model name in the request. LLM traffic differs from a typical API in three ways.
Requests vary a lot in cost. “Hi” and a 30,000-token document with a 2,000-token answer are both one request. A limit of N requests per minute doesn’t track what a user costs, so LLM gateways count tokens. Envoy AI Gateway reads the counts from the response: “AI Gateway automatically extracts token usage from LLM responses that follow the OpenAI schema format.” The count is only known once the answer is finished, so a request that pushes a user over the limit still completes; the user’s next request gets a 429.
Responses stream. Tokens go back as decode produces them (step 5 of part 1’s animation), so one request holds its connection open for the whole answer, which can be seconds to minutes. As far as I can tell, that means the gateway’s timeouts and connection draining during a deploy have to be set for minutes, where a typical API is tuned for requests that finish in milliseconds.
Some requests matter more than others. A chat user waiting on screen and an overnight batch job can share the same GPUs. The Kubernetes Gateway API Inference Extension lets a request carry a priority. An earlier version of its design spelled out what that meant: a Critical request goes to a server “with a wait queue lower than 50”, and a Sheddable request only goes to servers “that have lower than 80% KV cache utilization and fewer than 5 requests waiting in its queue”. The same extension adds an InferencePool to Kubernetes, “a set of Inference-focused Pods and an extension that will be used to route to them”, which the gateway uses in place of a Service. The extension is the Endpoint Picker, and it chooses the pod using “metrics emitted from the model servers”. That’s the router.
The router Link to heading
The CNCF post above lists “round robin, least request, ring hash, zone aware, priority” as load balancing algorithms that weren’t built for GPU-backed LLMs. There are two reasons.
First, counting requests doesn’t measure load. A server with two long prompts in flight can be busier than one with twenty short ones.
Second, where a request lands changes how much work it takes. If a worker already has the start of this prompt in its KV cache (the same system prompt, or the earlier turns of a conversation), it skips that part of prefill. Round robin, in llm-d’s words, will “scatter related requests across different pods”, and each pod computes the same prefix again.
Scoring workers Link to heading
A KV-aware router scores every worker on both things and picks the cheapest. NVIDIA Dynamo’s KV router uses, in its basic form:
cost = overlap_score_weight x prefill_blocks + decode_blocks
prefill_blocks blocks of this prompt the worker doesn't have cached and would have to compute
decode_blocks KV blocks the worker is already working on
Here’s a toy example with a prompt that fills 100 blocks and three workers. A has most of the prompt cached but is busy, B has none of it but is nearly idle, and C sits in between. Move the weight and watch which worker the request goes to.
At weights up to 2, B wins: computing the whole prompt on an idle worker is cheaper than joining a busy one. At 3, C wins. From 4 up, reusing A’s cache outweighs its load. A higher weight favors cache reuse and a faster first token. A lower weight spreads the load, which keeps each user’s token rate steady.
Knowing what’s cached Link to heading
The router needs to know what each worker holds, and there are two ways to find out.
The first is to ask the workers. Dynamo workers publish an event when they store new KV blocks and another when they evict some, and the router keeps a prefix tree of which worker holds which blocks.
The second is to guess from routing history. SGLang’s router “maintains an approximate radix tree of the actual radix tree on the workers” (SGLang v0.4): it records where it sent each prompt and assumes the prefix is still there. It needs nothing from the workers, but it can’t see evictions.
llm-d supports both, and benchmarked them on Qwen-32B, 8 vLLM pods of 2 H100s each, with 150 simulated customers sharing 6,000-token contexts:
routing p90 time to first token output tokens/s
precise (events) 0.54 s 8,730
approximate (history) 31 s 6,944
load only, no cache scoring 95 s 4,429
random 93 s 4,429
This is llm-d’s own benchmark on a workload built around shared prefixes, so the gap is at the large end. It still shows how much of the latency comes from routing and not from the GPUs.
Picking the decode worker Link to heading
For a disaggregated deployment the router also picks a decode worker, and there the cache overlap doesn’t matter, because the cache arrives from prefill. Dynamo turns overlap scoring off for decode “because decode routing should not chase prefix reuse”, and picks on the KV blocks already in use plus the ones this request will add. DeepSeek’s decode load balancer balances “KVCache usage across GPUs” and “request counts per GPU”.
Dynamo (experimentally) and llm-d can also prefer a decode worker in the same rack or zone as the chosen prefill worker. On an NVL72 that keeps the handoff on NVLink, as part 1 described.
When there’s too much traffic Link to heading
Disaggregation adds a way to waste work under load. Moonshot’s Mooncake paper (§7) describes it: “if a request is rejected by the decoding instance due to high load after the prefill stage has been completed, the computational resources expended during the prefill stage are wasted”. Their scheduler predicts what the decode load will be once a request’s prefill finishes, and rejects the request at the door if decode won’t have room.
The autoscaler Link to heading
The usual Kubernetes signal, CPU or GPU utilization, doesn’t work here. llm-d’s docs put it directly: “GPU utilization is often pegged near 100% during active batching regardless of actual load”. A server batching 4 requests and one batching 40 can both read 100%. Google’s GKE guidance recommends scaling on the server’s queue size, or its batch size for latency-sensitive services.
The signals that work come from the inference server: how many requests are waiting, how many are running, and how full the KV cache is. vLLM exports these as Prometheus metrics (vllm:num_requests_waiting, vllm:num_requests_running), and KServe can scale on them through KEDA.
With disaggregation, prefill and decode scale separately, since one depends on input tokens and the other on output tokens. Dynamo’s planner sizes each pool from a forecast of traffic and throughput per GPU measured in advance: prefill from input tokens per second, decode from output tokens per second. It then corrects both by how far the measured time to first token and time between tokens are from their targets.
Scaling also runs on a longer cycle. DeepSeek runs inference “across all nodes during peak daytime hours” and at night hands nodes to research and training; over their published day, inference peaked at 278 nodes and averaged 226.75.
All of this assumes a new replica starts serving soon after the autoscaler asks for it. The next section measures how long that takes.
Why a new replica takes minutes Link to heading
A new replica lands on one of two kinds of node. Either the node has served this model recently, so the container image is already there and the weights are on local disk or a shared volume, or it hasn’t: it’s new, back from repair, just updated, or moved over from training or another model. A large fleet sees both all the time. I timed the first case with vLLM 0.30 on Modal H100s, for Qwen2.5-7B on one GPU and Qwen2.5-32B on two, from vllm serve starting to the server answering /health. Every start ran in a fresh container, four times each; the table shows the median and the range. The script is in the lab repo.
7B 7B, caches kept 32B 32B, caches kept
python imports 28 s (23-33) 23 s (17-27) 29 s (17-52) 20 s (15-32)
engine and worker start 34 s (31-36) 18 s (17-20) 62 s (49-75) 50 s (35-63)
weight load 11 s (10-18) 6 s (5-9) 28 s (23-48) 10 s (8-17)
torch.compile 24 s (20-25) 2 s (1-2) 58 s (47-83) 3 s (2-3)
CUDA graph capture 12 s (11-13) 10 s (10-13) 104 s (86-214) 34 s (25-229)
FlashInfer JIT, KV setup 84 s (84-86) 6 s (3-8) 86 s (81-101) 6 s (5-7)
server up 4 s (3-6) 4 s (3-5) 7 s (6-10) 6 s (5-10)
time to /health 200 s (189-204) 68 s (68-75) 430 s (310-471) 139 s (97-343)
The medians of the phases don’t add up exactly to the median total, since each row’s median can come from a different run.
The first and third columns are what a pod gets by default. vLLM keeps its compile caches under ~/.cache inside the container, so they disappear with the pod, and the next pod on the same node compiles everything again. Compiling was most of the start: torch.compile, the kernels built during the 32B’s graph capture, and about 85 seconds in which the log goes quiet. During that quiet stretch the only files written were FlashInfer kernels built for the H100, in ~/.cache/flashinfer. vLLM uses FlashInfer for sampling, and on first use FlashInfer compiles its kernels with nvcc. That stretch took 81 to 101 seconds in every cold start, for both models.
The second and fourth columns are a second start with those caches kept, which is what a deployment gets when it mounts the cache directory on a volume or bakes it into the image. vLLM’s docs suggest copying the torch.compile cache between deployments “to save a great amount of compilation time”. Here keeping the caches saved over two minutes on the 7B and almost five on the 32B, and it’s a deployment setting, with no change to the model or the hardware. Even with every cache kept, Python imports and starting the engine took about 40 seconds on the 7B and 70 on the 32B.
One 2-GPU machine was slow in both of its starts: graph capture took 214 s cold and 229 s with caches, against 25 to 86 s elsewhere. That’s where the 32B’s wide ranges come from.
Weight loading grew with size: a median of 11 s for 14 GiB and 28 s for 65 GB from Modal’s volume. The weights were already on the volume; downloading the 32B from Hugging Face took another 143 s for 65.5 GB.
vLLM has a flag that skips compiling, --enforce-eager. On the 7B it brought the median start from 200 s to 166 s, because the FlashInfer kernels still get built, and it cut decode speed for one user from 166 to 72 tokens/s. Eager decode also varied from machine to machine, between 55 and 92 tokens/s, while compiled decode stayed between 162 and 168. That’s a bad trade for a server that will run for days.
A replica that answers /health is also slower than an old one for a while, because its prefix cache is empty. The same prompt of about 30,000 tokens took 1.03 s to its first token on the first request and 0.09 s when repeated on the same server. An old replica has the system prompts and recent conversations cached; a new one computes all of it, and a cache-aware router sends it fewer of the requests that would hit.
On a node that hasn’t served the model, more comes first: the image is pulled, the weights are fetched, and for a new node, the node boots and joins the cluster. I didn’t measure that case. The vllm/vllm-openai image is 8.73 GB compressed, and BentoML measured 5 to 6 minutes to pull it plus 3 to 4 to extract it. On AWS p5 nodes, loading 203 GiB of weights from S3 took about 423 s, 92% of that startup. Tools such as NVIDIA’s Run:ai Model Streamer and Modal’s GPU snapshots exist to cut these steps.
Caveats Link to heading
- The router, gateway and autoscaler numbers come from the projects’ own docs and benchmarks. I haven’t run Dynamo, llm-d or SGLang’s router.
- The Dynamo cost formula shown is the basic form from its docs; the current design adds terms for cache hits in host memory, disk and shared storage.
- The startup numbers are for two model sizes on Modal H100s. Modal loads container images lazily from its own store, which probably slows Python imports compared with a node that has the image fully pulled. Start times varied by 20% or more between runs of the same configuration.
- A production model of 70B or more, or a large MoE model split across nodes, spends much more of its startup on weights than these two did.
Next Link to heading
Part 3 is about hardware going bad: what a failed GPU, NIC or NVL72 tray does to the pools, the router and the requests in flight.