Part 3 asked whether a cache that already exists is cheaper to copy back to the GPU or to rebuild, and found copying wins: “any link faster than about 2 GB/s (16 Gb/s) beats rebuilding”. That was a script moving one cache by hand. This part looks at the same idea inside a serving engine, where it matters every day: a chat app resends the whole conversation on every turn, and an agent resends its context on every call. Inference in production part 5 quoted a measurement of Claude Code: “a median Claude step reads back 126k prefix tokens but appends only 857”. If the server still has the cache for those 126k tokens, it only has to prefill the 857 new ones.

Prefix caching Link to heading

When a request arrives, vLLM checks whether it already has the KV cache for the start of the prompt. The cache is stored in blocks of tokens, so a new request can reuse every block whose tokens match an earlier prompt from the start, and prefill only the rest. This is prefix caching, and vLLM turns it on by default.

Only a matching start counts. The cache for a token depends on every token before it, so a document reused after a different first sentence can’t reuse anything.

Setup Link to heading

  • Qwen3-8B in BF16 on one H100 80GB on Modal, vLLM 0.30 with prefix caching on.
  • Each prompt is a document of random tokens followed by a 32-token question, sent as token IDs so the lengths are exact. Each request asks for one output token, so its time is the time to first token, and only one request runs at a time.
  • A document’s first request is a miss. The same document with a different question is a hit: the document’s cache can be reused, and only the question needs prefill.
  • The script and raw results are in the lab repo.

Hits and misses Link to heading

Three documents at each length, from one run (a first run gave the same numbers to within a few milliseconds):

document    first request (miss)    same document again (hit)    faster
1k tokens          33 to 38 ms                13 to 14 ms         2.6x
4k tokens       106 to 108 ms                17 ms                6.3x
16k tokens      520 to 526 ms                30 ms               17.4x

A hit costs about the same at any length, because only the 32-token question is new. A miss grows with the document. So the longer the shared part, the more a hit saves: at 16k tokens, the user sees their first token after 30 ms instead of half a second.

Running out of room Link to heading

The cache doesn’t stay forever. vLLM had room for 388,560 tokens of cache on this GPU, and each document here takes 16,416 tokens with its question:

KV per token for Qwen3-8B          147,456 bytes
one 16k document                   16,416 x 147,456 bytes = 2.42 GB
documents that fit                 388,560 / 16,416       = about 23

To see what happens when it fills up, I read one 16k document, then 36 other 16k documents, enough to fill the cache one and a half times, then the first document again:

first document, first time          527 and 533 ms
first document, right away again     31 and 27 ms
36 other documents                  530 ms each (median)
first document, after the others    529 and 531 ms

By the time the first document came back, vLLM had dropped its cache to make room for the others, so it was a full miss again. On a busy server, an agent that pauses to run a tool can come back to find its cache gone.

Keeping evicted caches in CPU memory Link to heading

A GPU’s memory is small next to the machine it’s in. The H100 has 80 GB, and this container was given 192 GiB of the host’s CPU memory. LMCache plugs into vLLM and keeps a copy of each cache outside the GPU, so when vLLM evicts one, the copy can be loaded back instead of rebuilt. I gave it 100 GB of CPU memory and ran the same test:

                                        GPU only     LMCache, CPU memory
first document, first time              527, 533           609, 614 ms
first document, right away again         31,  27            30,  29 ms
36 other documents (median)             530, 530           610, 614 ms
first document, after the others        529, 531            69,  73 ms

After eviction, the first document came back in 69 and 73 ms instead of 530: 7.5 times faster than rebuilding. Loading the cache from CPU memory replaced the 16k-token prefill. As a speed:

2.42 GB / 69 ms  =  35 GB/s
2.42 GB / 73 ms  =  33 GB/s

That’s a lower bound, since the 69 ms also includes prefilling the question and returning a token. Part 3 measured “about 27 GB/s from pinned memory” on a different H100 machine, the same order.

Making the copies costs time. With LMCache on, every miss took about 610 ms instead of 525, 16% longer, while it also copied the new cache out to CPU memory. And it logged warnings near the end that it had run out of space for new blocks (the 37 documents need 89.6 GB, near its 100 GB), along with a few hundred internal warnings about pinned memory.

And on disk Link to heading

Disk holds far more than CPU memory, so I tried LMCache with a local disk too. This needed a CPU buffer as well: LMCache stages each cache in CPU memory on its way to disk, and with a 2 GB buffer it stalled waiting for space, logging “No eviction candidates found in local cpu backend”. With 20 GB, too little to hold the 36 other documents, the first document had to come back from disk:

                                        GPU only     LMCache, disk
first document, first time              527, 533       1,075, 1,108 ms
first document, right away again         31,  27         138,   152 ms
36 other documents (median)             530, 530       1,268, 1,269 ms
first document, after the others        529, 531         475,   426 ms

After eviction, the document came back in 475 and 426 ms, a little faster than rebuilding it, at 2.42 GB / 426 to 475 ms = 5.1 to 5.7 GB/s. That matches the disk speed part 3 measured on its own: “403 ms from local disk” for a 1.8 GiB cache. But writing every new cache to disk made each miss about twice as slow, and even a hit right after a miss took 138 ms instead of 30.

I can’t tell how much of the 426 to 475 ms was loading from disk. LMCache logged thousands of warnings, among them that it couldn’t allocate memory for a block while its CPU buffer was full, so some blocks may have been rebuilt on the way back instead of read. What the run does show is that for a 16k-token document on an 8B model, rebuilding takes about half a second on an H100, and disk didn’t beat that by much. Disk pays off when rebuilding costs more: longer documents, and larger models.

Caveats Link to heading

  • One H100, one setup, one request at a time. Under load, reloading from CPU memory competes with other requests for the same PCIe link, and the miss overhead adds to every new prompt.
  • Random tokens. Real prompts behave the same for prefix caching, which only compares token IDs, but the documents here never share a prefix by accident.
  • LMCache’s settings matter. The CPU size, chunk size and how it evicts decide what survives. These are one choice of settings, with LMCache 0.5.5.
  • Memory sizes. The container had 192 GiB of CPU memory, and LMCache was given 100 GB of it. Servers with more CPU memory can keep more caches.