<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
	<channel>
		<title>Prefix Caching on </title>
		<link>https://hiren.me/tags/prefix-caching/</link>
		<description>Recent content in Prefix Caching on </description>
		<generator>Hugo</generator>
		<language>en</language>
		
		
		
		
			<lastBuildDate>Thu, 08 Oct 2026 17:30:00 -0700</lastBuildDate>
		
			<atom:link href="https://hiren.me/tags/prefix-caching/index.xml" rel="self" type="application/rss+xml" />
			<item>
				<title>Watching a KV cache grow (part 4): reusing it</title>
				<link>https://hiren.me/posts/watching-a-kv-cache-grow-part-4/</link>
				<pubDate>Thu, 08 Oct 2026 17:30:00 -0700</pubDate>
				<guid>https://hiren.me/posts/watching-a-kv-cache-grow-part-4/</guid>
				<description>&lt;p&gt;&lt;a href=&#34;https://hiren.me/posts/watching-a-kv-cache-grow-part-3/&#34; &gt;Part 3&lt;/a&gt; asked whether a cache that already exists is cheaper to copy back to the GPU or to rebuild, and found copying wins: &amp;ldquo;any link faster than about 2 GB/s (16 Gb/s) beats rebuilding&amp;rdquo;. That was a script moving one cache by hand. This part looks at the same idea inside a serving engine, where it matters every day: a chat app resends the whole conversation on every turn, and an agent resends its context on every call. &lt;a href=&#34;https://hiren.me/posts/inference-in-production-part-5/&#34; &gt;Inference in production part 5&lt;/a&gt; quoted a measurement of Claude Code: &amp;ldquo;a median Claude step reads back 126k prefix tokens but appends only 857&amp;rdquo;. If the server still has the cache for those 126k tokens, it only has to prefill the 857 new ones.&lt;/p&gt;</description>
			</item>
	</channel>
</rss>
