<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>FP4 on </title>
    <link>https://hiren.me/tags/fp4/</link>
    <description>Recent content in FP4 on </description>
    <generator>Hugo</generator>
    <language>en</language>
    <lastBuildDate>Sun, 04 Oct 2026 12:00:00 -0700</lastBuildDate>
    <atom:link href="https://hiren.me/tags/fp4/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Weight precision (part 2): how FP4 works, and what it costs</title>
      <link>https://hiren.me/posts/weight-precision-part-2/</link>
      <pubDate>Sun, 04 Oct 2026 12:00:00 -0700</pubDate>
      <guid>https://hiren.me/posts/weight-precision-part-2/</guid>
      <description>&lt;p&gt;&lt;a href=&#34;https://hiren.me/posts/weight-precision-part-1/&#34; &gt;Part 1&lt;/a&gt; went down the ladder from FP32 to FP4 and what each step changes in production. This part stays on the last step. A weight stored in BF16 or FP16 can take any of 65,536 values. A weight stored in FP4 can take one of 16. Round Qwen3-8B&amp;rsquo;s weights to FP4 the obvious way, and all but 523 of the 50 million weights in one of its matrices become 0. Yet in one run, an FP4 version of the same model answered 1,205 of 1,319 grade-school math problems correctly, against 1,231 for the BF16 model.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Weight precision (part 1): how it affects serving</title>
      <link>https://hiren.me/posts/weight-precision-part-1/</link>
      <pubDate>Sun, 04 Oct 2026 11:00:00 -0700</pubDate>
      <guid>https://hiren.me/posts/weight-precision-part-1/</guid>
      <description>&lt;p&gt;FP32, BF16, FP8 and FP4 turn up all over LLM serving. A checkpoint is called &lt;code&gt;Qwen3-8B-FP8&lt;/code&gt; or &lt;code&gt;Qwen3-8B-NVFP4&lt;/code&gt;. vLLM takes flags like &lt;code&gt;--dtype bfloat16&lt;/code&gt; and &lt;code&gt;--kv-cache-dtype fp8&lt;/code&gt;. A GPU datasheet lists a different speed for BF16, FP8 and FP4. Each is a number format: it sets how many bits store every number in the model, from 32 in FP32 down to 4 in FP4. The choice decides how many GPUs a model needs, how many users fit, how fast tokens come out, and how accurate the answers are.&lt;/p&gt;</description>
    </item>
  </channel>
</rss>
