Part 1 followed a request through prefill and decode pools, and part 2 covered the gateway, router and autoscaler in front of them. This part is about what happens when a GPU, a node or an NVL72 tray fails under them: how often that happens, how it shows up, how much of the deployment it takes out, and what becomes of the requests that were running on it. The failure data comes from published papers and NVIDIA’s docs; the part about requests in flight I measured by killing vLLM replicas on Modal. The script is in the lab repo.

How often Link to heading

The best public numbers come from training. Meta’s Llama 3 paper reports 466 interruptions in 54 days on 16,384 H100s, 419 of them unexpected, which is about one every three hours. “Approximately 78% of the unexpected interruptions are attributed to confirmed hardware issues”, and GPUs were the largest share:

cause                      interruptions   share of unexpected
faulty GPU                 148             30.1%
GPU HBM3 memory             72             17.2%
software bug                54             12.9%
network switch or cable     35              8.4%
host maintenance            32              7.6%
GPU SRAM                    19              4.5%
NIC                          7              1.7%
silent data corruption       6              1.4%

NVLink has no row of its own. The paper describes its failures in prose: they “often manifest as stalled load/store operations within CUDA kernels without returning a clear error code.”

A study of NCSA’s Delta cluster (SC'25) found H100 memory “resilience is worse than A100”, with a 3.2x lower mean time between memory errors per GPU, and projects that “overprovisioning of 5% is necessary to handle GPU failures.”

I found no operator publishing failure rates for an inference fleet. The hardware is the same, so as far as I can tell the per-GPU rate carries over. What changes is the consequence. A training job spans every GPU, so one failure stops all 16,384; an inference fleet is many independent replicas, so one failure takes out the replicas on that hardware and the requests they were serving.

How a failure shows up Link to heading

Failures come in two kinds, and they need different detection.

Some are loud. The GPU drops off the PCIe bus, the driver logs an XID error, the engine process crashes, and its connections close. NVIDIA’s XID catalog lists a recovery action for each code:

XID   meaning                              action
48    double-bit ECC error                 reset the GPU
64    row remapping failure                reset the GPU
74    NVLink error                         follow NVIDIA's NVLink triage
79    GPU has fallen off the bus           reboot the host
95    uncontained memory error             reset the GPU
119   GSP RPC timeout                      reset the GPU

Some are quiet. A kernel stalls on NVLink, or the GPU stops making progress, and nothing errors. The process is still alive, its sockets are still open, and requests just stop moving.

I reproduced both on Modal with two vLLM replicas of Qwen2.5-7B, one per H100, and my own script as the client and router. Killing a replica’s processes with SIGKILL stood in for a loud failure; freezing them with SIGSTOP stood in for a quiet one. Neither is a real hardware fault, but both leave the serving stack in the state a real fault would: a worker that’s gone, or one that’s stuck.

When the replica died mid-answer, the client’s stream just ended: no exception, no error status, only a missing data: [DONE] at the end. My first version of the script treated that as a finished answer. A client or router that doesn’t check for the end marker will hand users a truncated answer as if it were complete. Checking for it, the client knew within 0.05 to 0.06 seconds.

The quiet failure gave nothing to check. The stream stopped after 128 tokens and the connection stayed open until my client’s 10-second read timeout fired. A /health request to the frozen replica didn’t fail either; it hung until its own 5-second timeout. How fast a hang is noticed depends entirely on those timeouts.

Kubernetes adds its own delays. NVIDIA’s device plugin marks a GPU unhealthy on most XIDs, but that only stops new pods from being scheduled on it: “Pods that were assigned to the failed devices will continue be assigned to this device”, says the Kubernetes documentation. Something else has to take the pod out of service. With the default readiness probe (every 10 seconds, three failures), the Gateway API Inference Extension’s endpoint picker drops a pod about 30 seconds after it goes bad, since it only routes to pods that are ready. If a whole node disappears, the node controller waits 50 seconds before marking it lost, and its pods stay bound for five more minutes. Tools like NVIDIA’s NVSentinel exist to watch DCGM and the kernel log and cordon and drain faster than that.

How much breaks: node or rack Link to heading

The NVLink boundary from part 1 decides how much of the deployment one failure takes out.

On HGX B200, a bad GPU takes out its node: one replica if the model spans all eight GPUs, or one of several smaller replicas. The rest of the fleet keeps serving.

On GB200 NVL72, it depends on which tray fails. From NVIDIA’s partition guide:

  • A compute tray: its four GPUs are marked NO_NVLINK, and the other 17 trays keep working. But “it causes errors on the partition workload”, so a model spread across the whole rack’s NVLink domain loses its replica.
  • A switch tray: in a single rack, “a switch tray/switch failure causes all GPUs in all partitions to lose NVLink connectivity.” Before repairing a switch tray, all partitions in the NVLink domain have to be deleted, so the repair takes the whole rack down too.

SemiAnalysis, an analyst firm, wrote that “the NVLink copper backplane still is not that reliable”, even after burn-in. I couldn’t find published failure rates for NVL72 racks.

This is the other side of part 1’s trade-off. NVL72 keeps the KV handoff on NVLink and makes 72-GPU replicas possible, and a 72-GPU replica is also what one failure can take out.

What happens to the requests in flight Link to heading

A request’s KV cache lives in the memory of the GPUs that were serving it. When they fail, the cache is gone, and the router has three choices: return an error, start the request again on another worker, or continue it there from the tokens already generated. NVIDIA Dynamo calls the third request migration. It’s “off by default”, and when on, “the migration system extracts the newly generated tokens and appends them to the request’s token sequence”. Only tokens move; the new worker computes the KV cache again from the prompt plus the answer so far.

I measured both recovery strategies. Replica A was killed after streaming 128 of 512 tokens, and replica B took over, either restarting the answer or continuing it. Prefix caching was off, so B never had the prompt cached from an earlier attempt.

For continue, my client sent B the prompt and the 128 tokens as token IDs, the way Dynamo does it.

prompt tokens   first new token on B     finish on B          continued answer
                restart    continue      restart   continue   matches uninterrupted
1,042           0.04 s     0.04 s        3.1 s     2.3 s      yes
8,692           0.23 s     0.22 s        3.3 s     2.6 s      yes
27,872          0.89 s     0.88 s        4.2 s     3.4 s      no

Both strategies pay the same price up front: B has to prefill the whole prompt again, which grows with the prompt from 0.04 seconds to almost 0.9. The 128 extra tokens that continue sends barely register. After that, continue wins. It generates only the rest of the answer, and the user doesn’t see 128 tokens a second time. With restart, the client has to throw away what it already showed, or show the user an answer that starts over.

Continuing matched the uninterrupted answer at 1k and 8.7k tokens but not at 28k. My guess is floating-point order: on B, the 128 tokens’ KV cache comes from one prefill, where the original run built it one decode step at a time, and at 28k tokens of context the small differences were enough to flip a later token. I haven’t confirmed that. In an earlier run I sent the answer so far as text, and it diverged at 8.7k as well, because re-tokenizing the text at the cut doesn’t always give back the same tokens.

Then I ran the same failure under load: eight users streaming from each replica, A killed six seconds in, and A’s eight users moved to B.

                                        restart       continue
moved users: gap before tokens resume   0.16-0.59 s   0.27-0.73 s
B's own users, time between tokens
  median, before the failure            6.5 ms        6.5 ms
  median, after                         7.4 ms        7.4 ms
  p99, before                           7.4-7.6 ms    7.2-7.4 ms
  p99, after                            8.4-8.6 ms    8.3-8.5 ms

The moved users were back within a second. B’s own users saw each token about a millisecond later, since B was now decoding 16 answers instead of 8, and the two strategies looked the same. With a 7B model and short prompts, one H100 had room for both replicas’ users; a replica already near its limit would have nowhere to put them, and the router would have to queue or reject the moved requests.

In the earlier text-based run, continue pushed B’s p99 to 36 ms. That run’s client sent each moved answer as text, which vLLM’s API server has to tokenize, and I suspect eight long tokenizations at once delayed the streaming that runs in the same process. I haven’t confirmed that either.

Disaggregation adds failure modes of its own. If the KV transfer from prefill to decode fails, vLLM’s kv_load_failure_policy decides what happens: fail (the default) returns an error, and recompute reschedules the request to compute the missing blocks again. If the decode worker dies before pulling the cache, the prefill worker keeps those blocks for kv_lease_duration, 30 seconds by default, before freeing them. Dynamo’s router returns a prefill or handoff failure to the request and doesn’t retry it another way.

Getting the capacity back Link to heading

The broken hardware has to be drained and diagnosed. DCGM’s deeper diagnostics need idle GPUs: level 2 takes up to 10.5 minutes on an 8-GPU system, level 3 up to 35. A GPU with a pending memory row remap needs a reset, which kills every process on it.

Meanwhile the lost capacity has to come back somewhere, and that’s part 2’s cold start. In my runs, restarting the killed replica on the same machine, with its compile caches still there, took 59 to 76 seconds across both runs, close to part 2’s cached number. On a node without the caches, part 2 measured over three minutes for a 7B. NVIDIA reports that Dynamo’s shadow engine, a standby process on the same node, restored capacity in 7.3 seconds where a cold restart took 283, for a GLM model on B200s. It covers a crashed process, not a lost GPU or node.

Caveats Link to heading

  • Killing or freezing a process isn’t a hardware fault. Real faults can be partial: a degraded NVLink, a flapping NIC, a GPU that’s slow but not stuck.
  • My failover numbers are for a 7B model, one GPU per replica, on one machine, with my own client as the router. Replicas split across 8 or 72 GPUs lose more at once and take longer to replace.
  • Each failover case ran once per strategy, so the numbers show the shape of the costs, not their spread.
  • The failure rates are from training fleets. I couldn’t find published rates for inference fleets or for GB200 NVL72 racks.

Next Link to heading

Part 4 covers rolling out a new model: getting new weights onto the fleet and moving traffic over without the cold starts and failures from the last two posts landing on users.