The first three parts followed a request through prefill and decode pools, the router and autoscaler in front of them, and what happens when the hardware under them fails. This part is about changing what runs on all of it: a new model, a new checkpoint of the same model, or a new engine version.
A rollout does to every replica, on purpose, what the last two posts measured happening by accident. Each old replica is shut down, like part 3’s failures, and each new one pays part 2’s cold start. A rollout goes well when users don’t notice either. The sources are project docs and operators’ own write-ups; the part about draining I checked on Modal.
Getting the weights there Link to heading
A new model has to reach every node before a replica can start. A 70B model in bf16 is 140 GB; DeepSeek-V4 Pro is 806 GiB. Pulling that from a model hub on every node is slow, and llm-d’s loading guide warns that “concurrent model weight downloads across Pods…can trigger Hugging Face rate limiting (HTTP 429)”. Fleets stage weights closer to the GPUs:
- Object storage, then local disk. On AWS p5 nodes, copying a 200 GB model from S3 to local NVMe took about 25 s, and streaming it onto the GPUs another 20. KServe’s LocalModelCache runs a download job on each node ahead of time.
- From another GPU. NVIDIA’s ModelExpress lets a new replica copy weights from a running one over NIXL and RDMA. For the 806 GiB model on 8xB200 nodes, the copy took under 10 seconds and startup went from 8 minutes to 1 minute 44. NVIDIA’s post adds that once weights are fast, “JIT kernel compilation (torch.compile, Triton, DeepGEMM, etc.)” is the largest cost left, which is what part 2 measured.
Moving traffic Link to heading
A Kubernetes rolling update replaces pods a few at a time. maxUnavailable and maxSurge both default to 25%: a quarter of the replicas can be down at once, or a quarter extra can be started first. Surge needs spare GPUs, which a GPU fleet often doesn’t have. Without surge, capacity drops for as long as each batch takes to come back. With part 2’s numbers, eight replicas of a 32B model rolled two at a time means the fleet runs at 75% for four waves of 2 to 7 minutes each, before counting weight downloads.
Rolling updates also give no way to test the new version on a small share of traffic first. The inference-aware gateways do:
- The Gateway API Inference Extension splits traffic between an old and a new InferencePool with weights on the route, for example 90/10, and Google’s guide says to “keep the original
InferencePooland nodes active during the roll out to enable rollbacks”. ItsInferenceModelRewriteresource does the same by model name inside one pool. - KServe has
canaryTrafficPercent, and setting it back to 0 rolls back. - Dynamo gives each version of its workers its own namespace, so old and new workers can’t find each other.
A canary pool also starts with empty prefix caches, and a KV-aware router will favor the old pool’s warm workers. As far as I can tell, that makes the new version’s first-token times look worse than they will be once its caches fill, which is worth knowing before rolling back on latency alone.
LoRA adapters are the cheap case. vLLM can load an adapter into a running server through /v1/load_lora_adapter, so rolling out a fine-tune moves megabytes onto pods that keep running. The vLLM docs warn that runtime loading “comes with security risks” and shouldn’t be used in production outside a fully trusted environment.
Draining the old version Link to heading
When a rolling update removes a pod, Kubernetes sends SIGTERM to its main process, waits the pod’s grace period (30 seconds unless set), and then sends SIGKILL. What the old replica does with the answers it’s still streaming decides whether users see the rollout.
According to llm-d’s shutdown guide, by default “vLLM immediately aborts in-flight requests and exits”. With --shutdown-timeout N, it keeps serving the running requests for up to N seconds first. I checked both on Modal with a vLLM 0.30 replica of Qwen2.5-7B, five users streaming 2,048-token answers, and a SIGTERM three seconds in. The script is in the lab repo.
default --shutdown-timeout 60
tokens each user got after SIGTERM 0 or 1 about 1,586 (all 2,048 in total)
answers completed 0 of 5 5 of 5
how the streams ended hung, then reset normally, with [DONE]
process exited after never (SIGKILL 13.6 s
at 90 s)
/health 0.5 s after SIGTERM connection answered
refused
With the default, vLLM stopped all five answers at once. Its log shows the engine shut down with “mode=abort timeout=0s”, and then the HTTP server printed “Waiting for connections to close” and waited. The five aborted streams were never closed, so the users’ connections stayed open with no tokens and no error until my stand-in for the grace period sent SIGKILL at 90 seconds. With Kubernetes’ 30-second default, that’s 30 seconds of a frozen answer, then a reset. This is vLLM 0.30; other versions may close the streams.
With --shutdown-timeout 60, the five answers finished in about 10 seconds and the process exited after 13.6. During that drain, /health kept answering. Kubernetes stops routing to a pod once it starts terminating, but a router that decides by polling /health would keep sending new requests to a replica that’s on its way out.
Part 3 showed why a cut-off answer is worse than an error: when the server closes a stream, the client sees the text end without data: [DONE], and a client that doesn’t check for the marker shows a truncated answer as if it were complete.
The timeout has to fit inside the grace period. llm-d says to set --shutdown-timeout “to roughly the p99 request duration” and terminationGracePeriodSeconds higher than that, with 90 and 120 seconds as its example. Dynamo will wait up to 900 seconds for in-flight work, but its pods default to a 60-second grace period, so Kubernetes kills them long before that wait runs out. OpenAI notes that “reasoning models can take several minutes to solve complex problems”. For those, a drain that lasts longer than any grace period is the only way to not cut answers, and the gap has to be covered by retrying the request elsewhere, like part 3’s request migration.
Prefill and decode on different versions Link to heading
Disaggregation adds a constraint the other parts didn’t have. A decode worker continues from the KV cache that a prefill worker computed, so both have to run the same weights. During a rollout, a new prefill worker and an old decode worker can both be up.
vLLM’s NIXL connector checks compatibility when the two connect. The hash it compares covers the vLLM version, the model’s architecture and shape, the attention backend and the KV cache dtype. It doesn’t cover the weights. An open pull request, #55776, says: “During a rolling update, prefill and decode replicas can therefore load different weights from the same model repository, pass the handshake, and mix KV states.” A fine-tuned checkpoint has the same shape as the one it replaces, so it passes, and the decode worker generates from a KV cache that different weights computed. Nothing errors; the answers are just built from two models.
The routers handle this above the engine. llm-d’s v0.9 release says that “during a model version rollout, the router understands which pods belong to which revision and routes accordingly, preventing the request failures that occurred when a prefill pod sent KV cache to a decode pod running a different model checkpoint.” Dynamo’s per-version namespaces keep new prefill workers with new decode workers, and an open design proposal, DEP #15214, covers keeping routers and workers of the same release together while both roll.
A bad model passes the health checks Link to heading
A replica can start, answer /health and stream tokens while giving worse answers, and none of the checks above will notice.
Anthropic’s postmortem of August and September 2025 describes three bugs of this kind. One was a misconfiguration deployed to its TPU servers that “caused an error during token generation”. Another was a routing error that a load-balancing change made worse, until it affected 16% of Sonnet 4 requests at its peak. Because routing was “sticky”, a user whose request hit a bad server was likely to keep hitting it. Their evaluations “simply didn’t capture the degradation”, and they said they “will run them continuously on true production systems”.
OpenAI rolled back a GPT-4o update in April 2025 after users found it too agreeable. Their write-up says their “offline evaluations…generally looked good”, and the A/B tests “seemed to indicate that the small number of users who tried the model liked it.” A rollback is a rollout too, with the same cold starts and drains.
The defenses are the ones this post already covered, used slowly: a small traffic share first, the old version kept running until the new one has served real traffic for a while, and evaluations that run against production and not only before launch. Gateway API can also mirror requests to a new pool without returning its answers, which lets a new version see production traffic before any user does. I didn’t find an operator describing how they compare the mirrored outputs.
Caveats Link to heading
- The drain check ran one 7B replica with five users. A replica serving hundreds of requests has more to drain and may need longer.
- The weight-distribution numbers are vendors’ own benchmarks.
- I haven’t run a canary or a rolling update of a real fleet. The capacity arithmetic uses part 2’s startup times and assumes nothing else goes wrong.
This is the last part I had planned. The series covered where the GPUs are, how a request finds them, what happens when they fail, and how to change what runs on them.