<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>FP8 on </title>
    <link>https://hiren.me/tags/fp8/</link>
    <description>Recent content in FP8 on </description>
    <generator>Hugo</generator>
    <language>en</language>
    <lastBuildDate>Sun, 04 Oct 2026 11:00:00 -0700</lastBuildDate>
    <atom:link href="https://hiren.me/tags/fp8/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Weight precision (part 1): how it affects serving</title>
      <link>https://hiren.me/posts/weight-precision-part-1/</link>
      <pubDate>Sun, 04 Oct 2026 11:00:00 -0700</pubDate>
      <guid>https://hiren.me/posts/weight-precision-part-1/</guid>
      <description>&lt;p&gt;FP32, BF16, FP8 and FP4 turn up all over LLM serving. A checkpoint is called &lt;code&gt;Qwen3-8B-FP8&lt;/code&gt; or &lt;code&gt;Qwen3-8B-NVFP4&lt;/code&gt;. vLLM takes flags like &lt;code&gt;--dtype bfloat16&lt;/code&gt; and &lt;code&gt;--kv-cache-dtype fp8&lt;/code&gt;. A GPU datasheet lists a different speed for BF16, FP8 and FP4. Each is a number format: it sets how many bits store every number in the model, from 32 in FP32 down to 4 in FP4. The choice decides how many GPUs a model needs, how many users fit, how fast tokens come out, and how accurate the answers are.&lt;/p&gt;</description>
    </item>
  </channel>
</rss>
