<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
	<channel>
		<title>Cost on </title>
		<link>https://hiren.me/tags/cost/</link>
		<description>Recent content in Cost on </description>
		<generator>Hugo</generator>
		<language>en</language>
		
		
		
		
			<lastBuildDate>Thu, 08 Oct 2026 17:01:00 -0700</lastBuildDate>
		
			<atom:link href="https://hiren.me/tags/cost/index.xml" rel="self" type="application/rss+xml" />
			<item>
				<title>Batching: how one GPU serves many users, and what a token costs</title>
				<link>https://hiren.me/posts/batching/</link>
				<pubDate>Thu, 08 Oct 2026 17:01:00 -0700</pubDate>
				<guid>https://hiren.me/posts/batching/</guid>
				<description>&lt;p&gt;A hosted model answers thousands of people at once, and a single user still gets their tokens quickly. Both come from batching, where one GPU works on many requests in the same step. This post explains why batching is nearly free at first, measures what happens as one H100 takes on more users, and turns the result into the cost of a token.&lt;/p&gt;&#xA;&lt;h2 id=&#34;why-batching-is-nearly-free-at-first&#34;&gt;&#xA;  Why batching is nearly free at first&#xA;  &lt;a class=&#34;heading-link&#34; href=&#34;#why-batching-is-nearly-free-at-first&#34;&gt;&#xA;    &lt;i class=&#34;fa-solid fa-link&#34; aria-hidden=&#34;true&#34; title=&#34;Link to heading&#34;&gt;&lt;/i&gt;&#xA;    &lt;span class=&#34;sr-only&#34;&gt;Link to heading&lt;/span&gt;&#xA;  &lt;/a&gt;&#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Generating a token reads the model&amp;rsquo;s weights from GPU memory. &lt;a href=&#34;https://hiren.me/posts/kernels/&#34; &gt;Kernels&lt;/a&gt; worked out what that costs for Qwen3-8B on an H100: each step reads 15.1 GB of weights, and the memory moves 3.35 TB/s, so a step takes at least 4.5 ms. The math in that step takes a small fraction of the time; the GPU mostly waits for the weights to arrive.&lt;/p&gt;</description>
			</item>
	</channel>
</rss>
