Part 1 went down the ladder from FP32 to FP4 and what each step changes in production. This part stays on the last step. A weight stored in BF16 or FP16 can take any of 65,536 values. A weight stored in FP4 can take one of 16. Round Qwen3-8B’s weights to FP4 the obvious way, and all but 523 of the 50 million weights in one of its matrices become 0. Yet in one run, an FP4 version of the same model answered 1,205 of 1,319 grade-school math problems correctly, against 1,231 for the BF16 model.

The reason to bother is bytes. Each decode step reads every weight, and in Watching a KV cache grow part 2 the time per step grew in a straight line with the bytes read. Qwen3-8B has 8.19 billion weights:

8.19 billion weights × 2 bytes   (BF16) = 16.4 GB
8.19 billion weights × 0.5 bytes (FP4)  =  4.1 GB

So FP4 should make the model 4 times smaller and its tokens up to 4 times faster. The FP4 checkpoint I used is 2.6 times smaller, and in one run on one B200, one user’s tokens came 1.54 times faster. This post works out, with the arithmetic at each step, how 16 values can hold a model, where the missing speed went, and what FP4 costs in accuracy.

What FP4 can hold Link to heading

The baseline in this post is Qwen3-8B as shipped, in BF16. FP4 uses 4 bits: 1 for the sign, 2 for the exponent and 1 for the mantissa. That gives 2^4 = 16 bit patterns. Here are all eight positive ones:

exponent bits   mantissa bit   value
00              0              0
00              1              0.5
01              0              1
01              1              1.5
10              0              2
10              1              3
11              0              4
11              1              6

The sign bit gives the negative copies. So every weight stored in FP4 has to be one of these, as NVIDIA’s description of the format lists them:

±0, ±0.5, ±1, ±1.5, ±2, ±3, ±4, ±6

Sixteen values aren’t enough Link to heading

Weights in a trained model are small numbers. These are the first four weights of one matrix in Qwen3-8B (layer 18’s down_proj), rounded to the nearest FP4 value:

weight     nearest FP4 value
 0.0132 →  0
-0.0471 →  0
-0.0109 →  0
-0.0347 →  0

Every weight becomes 0. The whole matrix goes the same way: its median weight is 0.017, and anything smaller than 0.25 rounds to 0. That’s all but 523 of its 50 million weights.

Block scaling Link to heading

The fix is to give each small block of weights its own scale. Divide the block by the scale, round to FP4, and multiply back when the weights are used. Here are the same four weights as one block:

1. Largest weight in the block:  0.0471
2. Scale = 0.0471 / 6 = 0.00785      (6 is the largest FP4 value)
3. Divide each weight by the scale:
      0.0132 / 0.00785 =  1.68
     -0.0471 / 0.00785 = -6.00
     -0.0109 / 0.00785 = -1.39
     -0.0347 / 0.00785 = -4.42
4. Round to the nearest FP4 value:   1.5,  -6,  -1.5,  -4
5. Multiply back by the scale:

   original    stored as          recovered    error
    0.0132      1.5 × 0.00785      0.0118       11%
   -0.0471     -6   × 0.00785     -0.0471        0%
   -0.0109     -1.5 × 0.00785     -0.0118        8%
   -0.0347     -4   × 0.00785     -0.0314       10%

Each weight is now off by 0 to 11%, where before every weight was lost. The block stores four 4-bit numbers plus one scale.

Why not one scale for the whole matrix? Because the largest weight sets the scale. This matrix’s largest weight is 1.21, about 70 times its median:

Scale = 1.21 / 6 = 0.2017

    0.0132 / 0.2017 =  0.07  →  0
   -0.0471 / 0.2017 = -0.23  →  0
   -0.0109 / 0.2017 = -0.05  →  0
   -0.0347 / 0.2017 = -0.17  →  0

All four are lost again. With small blocks, one large weight only spoils its own few neighbors.

The two common FP4 formats both use blocks. MXFP4 gives each block of 32 weights one 8-bit scale, which has to be a power of two. NVFP4 gives each block of 16 weights one 8-bit FP8 scale, plus one 32-bit scale for the whole matrix. The scales cost a few bits per weight:

MXFP4: 4 bits + 8 bits / 32 weights = 4.25 bits per weight
NVFP4: 4 bits + 8 bits / 16 weights = 4.5 bits per weight

I rounded the full layer 18 down_proj matrix all four ways on my laptop. “Average error” is the average distance between each weight and its stored value, divided by the average size of a weight.

                          weights that become 0   average error
straight to FP4           99.999%                 99.99%
one scale per matrix      94%                     93%
MXFP4 (32 per block)      9.3%                    11.1%
NVFP4 (16 per block)      7.3%                    9.0%

The same shard of the model has 70 more matrices, from layers 17 to 27. NVFP4’s error was between 9.0% and 9.2% in every one of them, and MXFP4’s between 10.8% and 11.5%. With one scale per matrix, the error ran from 17% to 95%, depending on how large that matrix’s biggest weight was. NVFP4 does a little better than MXFP4 because its blocks are half the size, and its scale can be any FP8 number, not only a power of two. The script is kv_fp4_error.py.

Where the missing speed went Link to heading

The FP4 checkpoint I used is RedHatAI/Qwen3-8B-NVFP4. It keeps two parts of the model in 16 bits: the embedding table that turns tokens into vectors, and the output layer that scores the next token. Each is 151,936 tokens × 4,096 numbers. From the checkpoint’s own file sizes:

FP4 weights:        6.95 billion × 0.5 bytes     = 3.47 GB
FP8 scales:         6.95 billion / 16 × 1 byte   = 0.43 GB
kept in 16 bits:    1.24 billion × 2 bytes       = 2.49 GB
total                                            = 6.40 GB   (BF16 checkpoint: 16.4 GB, so 2.6x smaller)

For speed, what matters is the bytes read on every decode step. The embedding table barely counts, because each step looks up only one row of it. With one user, each step reads the rest once, so the best case on a B200 is:

time per token ≈ bytes read per step / memory bandwidth

  BF16: (16.4 - 1.24) GB = 15.1 GB / 7.7 TB/s = 1.97 ms per token
  FP4:  ( 6.4 - 1.24) GB =  5.2 GB / 7.7 TB/s = 0.67 ms per token   → 2.9x faster, at best

Part 1 measured 4.07 ms for BF16 and 2.64 ms for FP4, in one run on one B200: 1.54x. (An earlier run, with kv_fp4_serve.py, gave the same numbers within 0.01 ms.) The gap makes sense if each step also spends a fixed time on work that doesn’t depend on weight bytes. Write each step as bytes / bandwidth + fixed time, and the two measurements give two equations:

BF16: 4.07 ms = 15.1 GB / bandwidth + fixed
FP4:  2.64 ms =  5.2 GB / bandwidth + fixed

subtract:  1.43 ms = 9.9 GB / bandwidth  →  bandwidth = 6.9 TB/s
then:      fixed = 4.07 ms - 15.1 GB / 6.9 TB/s = 1.9 ms

So in this run, reading the weights went at roughly 6.9 TB/s, close to the 7.7 TB/s on NVIDIA’s B200 datasheet. Two points make a rough estimate, not a measurement of bandwidth. But each step also spent about 1.9 ms on everything else: attention, sampling, launching kernels, the scheduler. FP4 doesn’t shrink that part, and at 4 ms per step it is almost half. This split assumes the fixed time is the same for both, which is only roughly true, because FP4 also has to scale its activations on every step.

With many users, the speedup shrinks further. FP4 makes the weights smaller, but the KV cache stays in 16 bits, and every step reads the cache too. Qwen3-8B has 36 layers, 8 KV heads and 128 numbers per head:

KV per token = 2 (K and V) × 36 layers × 8 heads × 128 numbers × 2 bytes = 147,456 bytes
16 users × 8,192 tokens = 131,072 tokens
131,072 × 147,456 bytes = 19.3 GB of cache, read on every step

BF16: 15.1 GB + 19.3 GB = 34.5 GB per step
FP4:   5.2 GB + 19.3 GB = 24.5 GB per step   → 1.4x faster

Part 1 measured 1.36x for 16 users at 8k, close to the bytes this time, because the cache dominates every step.

What FP4 costs Link to heading

I asked both models all 1,319 questions in the GSM8K test set (grade-school math word problems), with thinking off and greedy decoding, and checked the final number. This run, with kv_fp4_serve.py, saved every answer, so it shows which questions changed:

BF16:    1,231 correct   (93.3%)
FP4:     1,205 correct   (91.4%)

both right     1,175
BF16 only         56
FP4 only          30
neither           58

FP4 gave a different final answer on 112 questions, 8.5% of them. It lost 56 that BF16 got right, and it got 30 right that BF16 missed. Net, it lost 26 questions, about 2 points, and kept 98% of the BF16 score. A second run, on another B200 for part 1, got 1,233 for BF16 and 1,205 for FP4 again. The gap held, though individual answers may have moved.

RedHat’s own evals of this checkpoint show where the cost grows. On GSM8K, in their setup, it scored 86.73 against 87.26 for the BF16 model, keeping 99.4%. On harder tests it kept less: 79% on MMLU-Pro (27.49 against 34.64) and 82% on the AIME 2024 math problems (62.07 against 75.86). The 9% error per weight from the rounding test above barely shows on easy questions and shows clearly on hard ones.

Caveats Link to heading

  • One run of one model on one GPU, with one FP4 checkpoint.
  • Only Blackwell GPUs like the B200 do math in FP4 directly. Part 1 ran this checkpoint on an H100, where vLLM fell back to weight-only FP4: faster for one user, 1.75x slower on an 8k prompt.
  • This checkpoint also rounds activations to FP4 during the math, not only the weights. The rounding test covers the weights only.
  • GSM8K is one easy eval, run twice per model across two posts. The BF16 score moved by 2 questions between runs, so gaps that small are noise.
  • The speed tests used random prompts and forced 512-token answers, not real traffic.