This series is about how LLM inference is deployed in production: what the hardware looks like, where each piece of work runs, and how one request moves through it. The numbers in this post come from NVIDIA’s documentation and from what DeepSeek and Moonshot have published about their own systems. Where a piece overlaps with something I measured earlier, I link to it.

The short version: production systems split each request’s two phases across separate pools of GPUs, and the design question that changes the picture most is where the NVLink boundary sits. On an HGX B200 system it’s the node, so the KV cache crosses the network between the pools. On a GB200 NVL72 it’s the whole rack, so the cache can stay on NVLink, as long as the deployment fits inside that rack.

Why split at all Link to heading

A request does two different jobs. Prefill processes the whole prompt at once and keeps the GPU busy computing; decode then produces the answer one token at a time, reading the whole KV cache on every step. When both share a GPU, a long prompt stalls everyone who is already decoding there. I measured that in Inference disaggregation part 1: a 32k-token prompt froze another user’s output for 1.5 seconds.

Disaggregated inference gives each job its own pool. The prefill pool builds the KV cache and hands it to the decode pool, which generates the tokens. That handoff is the new piece of infrastructure, and most of this post follows it.

What a real deployment looks like Link to heading

Two teams have published enough detail to anchor this.

DeepSeek described their V3/R1 inference system in early 2025. They run separate prefill and decode deployments on H800 nodes (8 GPUs each), with different parallelism for each:

                 deployment unit       parallelism           per-node throughput
prefill          4 nodes (32 GPUs)     experts spread x32    ~73.7k tokens/s in
decode           18 nodes (144 GPUs)   experts spread x144   ~14.8k tokens/s out

peak nodes in use over 24 hours:   278   (average 226.75)
input tokens served from cache:    56.3% (from an on-disk KV cache)

The model is a large mixture-of-experts (MoE) model, and that’s the usual reason teams split. Prefill and decode want the experts spread across GPUs in different ways, and a single shared pool can only pick one.

Moonshot’s Mooncake (FAST ‘25, which serves Kimi) takes the idea further. It separates prefill and decode clusters, keeps KV caches in a pool built from the machines’ CPU memory, SSDs and NICs, and uses a global scheduler, which they call Conductor, to pick a prefill and a decode instance for each request. The paper reports it running across thousands of nodes and processing over 100 billion tokens a day.

Following one request Link to heading

Here’s the path a request takes through a disaggregated deployment. The dots show what’s moving (blue for the prompt, orange for the KV cache, green for generated tokens), and the strip underneath shows the cache itself for a five-token prompt: empty until prefill fills all five slots at once, then growing by one slot per generated token. Switch between the two systems at step 4, where they differ.

Following one request through a disaggregated deployment
System:
prompt
KV cache
user sees
the request (prompt) the KV cache generated tokens
Simplified: each pool is drawn as one node or a few trays, and real pools span many nodes. Step 4 on B200 draws a cache split by attention heads, one slice per GPU; when each request's cache sits on one GPU instead, its handoff uses a single rail. The cache strip uses 57 KB per token, measured for Qwen2.5-7B in the KV cache series (other models differ), and example tokens from a real SmolLM2 run.

The order matches how the serving software describes itself. NVIDIA’s Dynamo docs put it in one line: “Dynamo routes each request through prefill first, transfers the KV cache to the decode worker, then streams the response from decode.”

A KV-aware router also looks at caches: it prefers a prefill worker that already holds part of the prompt in its cache, so those tokens don’t have to be computed again. Reusing cached prompts matters at this scale: DeepSeek served 56.3% of its input tokens from an on-disk KV cache. The gateway and router get their own post next.

The building block: a rail-optimized fabric Link to heading

Before the two systems, one piece of networking they both use. In NVIDIA’s reference designs, every GPU gets its own 400 Gb/s InfiniBand NIC, and the NICs are wired in rails: every node’s first NIC goes to one leaf switch, every node’s second NIC to another, and so on.

A rail-optimized fabric: GPU N of every node plugs into leaf switch N

NVIDIA’s DGX SuperPOD reference architecture describes the result for a group of 32 B200 nodes: “Traffic per rail of the DGX B200 systems is always one hop away from the other 31 nodes in a SU. Traffic between nodes, or between rails, traverses the spine layer.” A group like that uses 8 leaf switches, one per rail, and 4 spines.

Where a request’s KV cache lives depends on how the model is split across GPUs. As I understand it:

  • Split by attention heads (tensor parallelism): every GPU holds some of the heads’ keys and values for every token, so each GPU has a slice of every request’s cache. If prefill and decode use the same GPU layout, each prefill GPU’s slice goes to the matching decode GPU on the same rail, one switch hop away, and all eight rails carry slices at once.
  • Split by request (data-parallel attention, which is how DeepSeek runs attention): each request’s whole cache sits on one GPU, so its handoff is one GPU to one GPU on a single rail. Different requests use different rails at the same time.

Either way, keeping the GPU layouts matched keeps the handoff on its rail and off the spines.

From NVIDIA’s DGX B200 user guide:

GPUs                 8x B200, 1,440 GB of GPU memory in total
inside the node      NVLink 5 through 2 NVSwitch chips, 1.8 TB/s per GPU
compute network      8x ConnectX-7, up to 400 Gb/s each, one per GPU
front-end/storage    2x BlueField-3
power                14.3 kW max

Prefill and decode pools are sets of these nodes, so every KV handoff leaves a node: GPU, NIC, leaf switch, NIC, GPU. With GPUDirect RDMA the data goes straight between GPU memory and the NIC, without passing through the CPU’s memory.

The gap between the two links is large. In Inference disaggregation part 2 I measured a GPU-to-GPU copy over NVLink on H100s at 395 GB/s. A 400 Gb/s NIC carries at most 50 GB/s in each direction. Within a node the cache moves over NVLink; between nodes, it goes through a NIC at about an eighth of that speed. When the cache is split by heads, all eight NICs share the work for each request; when each request’s cache sits on one GPU, that request’s handoff gets one NIC. The offload post in the KV cache series estimated what that costs for one 32k-token cache: roughly 3% on top of prefill at 400 Gb/s.

From NVIDIA’s DGX GB200 user guide and GB200 NVL72 page:

compute trays        18, each with 2 Grace CPUs and 4 Blackwell GPUs
NVLink switch trays  9, with 2 NVLink switch chips each
NVLink domain        all 72 GPUs, 1.8 TB/s per GPU, 130 TB/s across the rack
GPU memory           13.4 TB HBM3E in the rack
compute network      4x ConnectX-7 400G per tray, one per GPU
power                about 120 kW per rack (DGX GB200 figure), liquid cooled

Here the NVLink domain covers the whole rack. If a prefill pool and a decode pool both live inside one rack, the KV handoff goes through the NVLink switch trays and never touches a NIC. Each GPU still has its own NIC and the rails still exist, but they connect racks to each other.

There’s a limit. DeepSeek’s decode unit is 144 GPUs, twice the size of one NVL72 rack. At that scale the decode pool spans racks, and the fabric between racks carries KV traffic again. NVL72 keeps the handoff on NVLink when a deployment unit fits in the rack. DeepSeek’s 32-GPU prefill unit would; their decode unit wouldn’t.

Side by side Link to heading

HGX B200 GB200 NVL72
NVLink boundary the node, 8 GPUs the rack, 72 GPUs
Largest group of GPUs on NVLink 8 72
Prefill-to-decode handoff through the InfiniBand rails NVLink inside the rack, if both pools fit
What the rail fabric carries every KV handoff traffic between racks
Unit you add or replace a node a rack, or a tray within it
Power 14.3 kW per node about 120 kW per rack

The trade-off: NVL72 makes the handoff and wide expert parallelism easier, because 72 GPUs share one fast domain. The cost is a much bigger failure and scheduling unit. As far as I know, a failed tray takes its four GPUs out of the NVLink domain and disturbs any job spread across it, while on HGX a failure stays inside one node. HGX pays for that containment with network traffic on every disaggregated request.

The software that does this today Link to heading

  • NVIDIA Dynamo sits above an existing engine (vLLM, SGLang or TensorRT-LLM) and adds the front end, KV-aware routing, an autoscaling planner and cache management across GPU, CPU and SSD.
  • vLLM supports disaggregated prefill, still marked experimental, with a NIXL-based connector for the KV transfer. Its docs say plainly that it “DOES NOT improve throughput”: the gain is steadier latency.
  • SGLang supports prefill/decode disaggregation with Mooncake or NIXL as the transfer engine.
  • llm-d runs vLLM on Kubernetes with the pieces above as Kubernetes components, using the Gateway API Inference Extension to route requests.
  • NIXL, NVIDIA’s Inference Xfer Library, is the transfer layer several of these share. It picks NVLink, RDMA or storage underneath depending on where the cache is going.

Caveats Link to heading

  • I haven’t run anything on B200 or NVL72. Their figures are NVIDIA’s, and the deployment numbers are what DeepSeek and Moonshot published.
  • NVIDIA quotes NVLink bandwidth as 1.8 TB/s per GPU. As far as I can tell that counts both directions, while the 400 Gb/s NIC figure is per direction, so compare them with care.
  • DeepSeek’s numbers are for H800 nodes, not B200, and come from one 24-hour window they chose to publish.
  • The request-flow animation simplifies a lot: real pools span many nodes, and the prefill/decode split drawn inside the NVL72 rack is illustrative.

Next Link to heading

Part 2 covers the front door: the gateway, router and autoscaler that sit in front of these pools, how a request picks its workers, and why a new replica takes minutes to become useful.