<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
	<channel>
		<title>Kernels on </title>
		<link>https://hiren.me/tags/kernels/</link>
		<description>Recent content in Kernels on </description>
		<generator>Hugo</generator>
		<language>en</language>
		
		
		
		
			<lastBuildDate>Wed, 07 Oct 2026 17:52:48 -0700</lastBuildDate>
		
			<atom:link href="https://hiren.me/tags/kernels/index.xml" rel="self" type="application/rss+xml" />
			<item>
				<title>Kernels: what they are and why inference cares</title>
				<link>https://hiren.me/posts/kernels/</link>
				<pubDate>Wed, 07 Oct 2026 17:52:48 -0700</pubDate>
				<guid>https://hiren.me/posts/kernels/</guid>
				<description>&lt;p&gt;An LLM running on a GPU spends its time in kernels, the small programs the GPU runs one after another. In the measurement later in this post, generating one token with Qwen3-8B on an H100 took 2,398 of them. This post starts from the beginning: what a kernel is, how model code turns into kernels, where kernels come from, and which ones run for each token.&lt;/p&gt;&#xA;&lt;h2 id=&#34;why-a-gpu&#34;&gt;&#xA;  Why a GPU&#xA;  &lt;a class=&#34;heading-link&#34; href=&#34;#why-a-gpu&#34;&gt;&#xA;    &lt;i class=&#34;fa-solid fa-link&#34; aria-hidden=&#34;true&#34; title=&#34;Link to heading&#34;&gt;&lt;/i&gt;&#xA;    &lt;span class=&#34;sr-only&#34;&gt;Link to heading&lt;/span&gt;&#xA;  &lt;/a&gt;&#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Generating one token means multiplying a vector by every weight matrix in the model. Each output number is a long sum of multiplications, and every output number can be worked out at the same time as the others. A CPU has a handful of powerful cores and does a few of these at a time. A GPU has thousands of simpler ones and does many at once.&lt;/p&gt;</description>
			</item>
	</channel>
</rss>
