<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
	<channel>
		<title>Speculative Decoding on </title>
		<link>https://hiren.me/tags/speculative-decoding/</link>
		<description>Recent content in Speculative Decoding on </description>
		<generator>Hugo</generator>
		<language>en</language>
		
		
		
		
			<lastBuildDate>Thu, 08 Oct 2026 17:10:00 -0700</lastBuildDate>
		
			<atom:link href="https://hiren.me/tags/speculative-decoding/index.xml" rel="self" type="application/rss+xml" />
			<item>
				<title>Speculative decoding: guessing tokens to use idle compute</title>
				<link>https://hiren.me/posts/speculative-decoding/</link>
				<pubDate>Thu, 08 Oct 2026 17:10:00 -0700</pubDate>
				<guid>https://hiren.me/posts/speculative-decoding/</guid>
				<description>&lt;p&gt;With one user, a GPU spends most of each decode step waiting for the weights to arrive from memory, and its math units sit mostly idle. &lt;a href=&#34;https://hiren.me/posts/batching/&#34; &gt;Batching&lt;/a&gt; fills that idle math with other users&amp;rsquo; tokens. Speculative decoding fills it with guesses: something cheap guesses the next few tokens, and the model checks all of them in one step. When the guesses are right, one step produces several tokens. This post measures how often the guesses are right, and what that does to speed on an H100.&lt;/p&gt;</description>
			</item>
	</channel>
</rss>
