FP32, BF16, FP8 and FP4 turn up all over LLM serving. A checkpoint is called Qwen3-8B-FP8 or Qwen3-8B-NVFP4. vLLM takes flags like --dtype bfloat16 and --kv-cache-dtype fp8. A GPU datasheet lists a different speed for BF16, FP8 and FP4. Each is a number format: it sets how many bits store every number in the model, from 32 in FP32 down to 4 in FP4. The choice decides how many GPUs a model needs, how many users fit, how fast tokens come out, and how accurate the answers are.

This post explains what the formats are, then goes through what each one changes in production. For each point, it shows the arithmetic, which holds anywhere, and then one run of Qwen3-8B on a B200 and an H100, which holds only for that setup. Part 2 goes deep on FP4, the smallest of them.

What a number format is Link to heading

A model is billions of numbers, called weights. Each one is stored as a floating-point number, which works like scientific notation in base 2: a sign, an exponent and a mantissa. Here is 0.15625 in FP16, the 16-bit format:

0.15625 = 1.25 × 2^-3

FP16 bits:   0      01100      0100000000
             sign   exponent   mantissa

sign:      0                    → positive
exponent:  01100 = 12, minus a fixed offset of 15 = -3   → 2^-3 = 0.125
mantissa:  0100000000           → 1 + 1/4 = 1.25   (the bits are worth 1/2, 1/4, 1/8, ...)
value:     1.25 × 0.125 = 0.15625

The exponent bits decide how large or small a number can get. The mantissa bits decide how finely it can be set between two powers of two. Every format below makes a different trade between the two.

The ladder Link to heading

Each step down halves the bits, so each step stores a number in half the bytes, and with fewer values to choose from. The limits are part of each format’s definition, so they are the same in every framework and on every GPU. “Gap near 1” is the distance from 1 to the next number the format can hold.

format       sign/exponent/mantissa   bytes   largest value   gap near 1
FP32         1 / 8 / 23               4       3.4 × 10^38     0.00000012
FP16         1 / 5 / 10               2       65,504          0.00098
BF16         1 / 8 / 7                2       3.4 × 10^38     0.0078
FP8 (E4M3)   1 / 4 / 3                1       448             0.125
FP4 (E2M1)   1 / 2 / 1                0.5     6               0.5

Here is what happens to two numbers on the way down:

            3.14159       70,000
FP32        3.1415901     70,000
FP16        3.140625      too large (the largest is 65,504)
BF16        3.140625      70,144
FP8         3.25          too large (the largest is 448)
FP4         3             too large (the largest is 6)

FP16 and BF16 are both 2 bytes, but they split the bits differently. FP16 spends more bits on the mantissa, so it is more precise, but nothing above 65,504 fits. BF16 keeps FP32’s 8 exponent bits, so it covers the same range as FP32 with less precision. Training with FP16 needed extra tricks to keep values in range. Google’s write-up on BF16 puts it as “Unlike FP16, which typically requires special handling via techniques such as loss scaling, BF16 comes close to being a drop-in replacement for FP32”. Most open models are now trained and shipped in BF16. Qwen3-8B’s config says "torch_dtype": "bfloat16", and so does Mistral-7B’s.

FP8 and FP4 can’t hold weights directly. A largest value of 448, or 6, with steps of 0.125 or 0.5, is far too coarse for weights that are mostly around 0.01 to 0.1. So an FP8 or FP4 checkpoint also stores scale factors: a group of weights shares one scale, and each weight is stored relative to it. Part 2 works through how that works for FP4. FP8 also has a second layout, E5M2, with more range (up to 57,344) and less precision. The paper that defined both recommends “E4M3 for weight and activation tensors, and E5M2 for gradient tensors”, so serving uses E4M3.

What changes in production Link to heading

Each section starts with the arithmetic, which holds for any model. Then it shows one run: Qwen3-8B as shipped in BF16, Qwen’s own Qwen3-8B-FP8, and RedHat’s FP4 build, Qwen3-8B-NVFP4. Each was served by vLLM 0.30 on one Modal B200 and one H100, and each speed test ran twice. The script and raw results are in kv_formats.py.

Memory: how many GPUs Link to heading

This part is arithmetic. A model’s weights take parameters × bytes per weight. For a 70B model:

FP32:  70 billion × 4 bytes   = 280 GB
BF16:  70 billion × 2 bytes   = 140 GB   → at least 2 H100s (80 GB each), before any KV cache
FP8:   70 billion × 1 byte    =  70 GB   → fits on 1 H100, with only 10 GB left for the cache
FP4:   70 billion × 0.5 bytes =  35 GB   → plus a few GB of scales

Real checkpoints come out a little larger than this, because they keep some layers in 16 bits. The download shrinks with the weights, and so does the time a new replica spends fetching them at startup, which Inference in production part 2 measured.

Measured, as vLLM reported it at startup:

                     BF16        FP8         FP4
weight memory        16.4 GB     9.5 GB      6.4 GB
smaller than BF16    -           1.7x        2.6x

Neither reaches its 2x or 4x, for the same reason: both checkpoints keep the embedding table and output layer in BF16.

KV cache room: how many users Link to heading

Whatever memory the weights don’t use goes to the KV cache, which holds each request’s context. The KV cache has its own format, set separately from the weights with --kv-cache-dtype. For Qwen3-8B, the arithmetic is:

KV per token = 2 (K and V) × 36 layers × 8 heads × 128 numbers × bytes per number
  BF16 cache: × 2 bytes = 147,456 bytes per token
  FP8 cache:  × 1 byte  =  73,728 bytes per token   → twice the tokens in the same memory

Measured, in the same run: how many tokens of cache vLLM had room for.

tokens of KV cache                B200 (180 GB)        H100 (80 GB)
BF16 weights, BF16 cache          1,056,992            388,560
FP8 weights                       1,100,864   (+4%)    442,848   (+14%)
FP4 weights                       1,117,072   (+6%)    464,096   (+19%)
BF16 weights, FP8 cache           2,113,344   (2.0x)   not run

Smaller weights free the same 7 to 10 GB on both GPUs. On the B200 that’s a small share of the cache, and on the H100 a bigger one. The FP8 cache doubled the room, exactly as the arithmetic said. This run used vLLM’s default FP8 cache scales. Its docs recommend calibrating them.

Decode: limited by bytes Link to heading

Each decode step makes one token per request and reads every weight to do it. With few requests, the time per token follows the bytes read, so halving the bytes can at most halve the time. The embedding table (1.24 GB here) barely counts, because each step reads only one row of it. That gives a best case for 1 user, and the measured times, with a range where the two runs differed:

                           BF16          FP8           FP4
bytes read per step        15.1 GB       8.3 GB        5.2 GB
best case vs BF16          -             1.8x          2.9x

ms per token
B200, 1 user               4.07          3.47 (1.17x)  2.64 (1.54x)
H100, 1 user               6.60          4.51 (1.46x)  3.99 (1.65x)
B200, 16 users, 8k         8.20-8.26     7.19-7.22     6.05-6.08
H100, 16 users, 8k         17.13         13.89-13.91   17.94-18.15

Neither GPU reached the best case, because each step also has fixed costs that don’t shrink with the weights. The H100 came closer because its memory is slower. Reading 15.1 GB at 3.35 TB/s takes 4.5 ms of its 6.6 ms step. On the B200 it takes 2.0 ms of a 4.1 ms step, and FP8 or FP4 can’t shrink the other half.

With 16 users at 8k, the FP8 KV cache helped too: 6.82 ms per token against 8.23, because each step reads half the cache bytes. FP4 on the H100 got slower under load, which the hardware section explains.

Prefill: limited by math Link to heading

Prefill reads the whole prompt at once: thousands of tokens through every weight in one pass. That is limited by how many multiply-adds the GPU can do per second, and GPUs do more of them in smaller formats. Here are NVIDIA’s peak numbers for dense math. The H100 and B200 pages print figures “with sparsity”, a special case most models don’t use. NVIDIA’s B200 datasheet says “Dense is one-half of the sparse spec shown”, so these are halved:

TFLOPS, dense   BF16      FP8       FP4
A100            312       none      none
H100            989       1,979     none
B200            2,250     4,500     9,000

Sources: the A100 datasheet, the H100 page and the B200 datasheet. Each step down doubles the peak, on the GPUs that support it.

A checkpoint gets this speedup only if the math itself runs in the small format, not just the storage. Both checkpoints here do. Qwen3-8B-FP8’s config sets "activation_scheme": "dynamic", and the NVFP4 checkpoint quantizes activations too, so the multiply-adds run in FP8 or FP4.

Measured: time to first token for one 8,192-token prompt, so mostly prefill.

ms to first token, 8k prompt   BF16        FP8         FP4
B200                           136         106-125     80-84
H100                           249-252     196-201     433-442

On the B200, FP8 cut the time by 1.1 to 1.3x and FP4 by 1.6 to 1.7x. That’s less than the datasheet’s 2x per step, likely because parts of prefill, like attention, still run in 16 bits. On the H100, FP8 cut it by 1.25x. FP4 made it 1.75x slower.

Hardware support: not every GPU has every format Link to heading

The table above has gaps because the math units, called Tensor Cores, are built into the chip. Each GPU generation added formats, and an older chip can’t gain one later. The H100’s Tensor Cores added FP8 to what the A100 had, and NVIDIA’s Blackwell generation, which includes the B200, added NVFP4. So an A100 has no FP8 math, and an H100 has no FP4 math.

A checkpoint can still load on a GPU that lacks its format. Then vLLM falls back to keeping the small weights in memory and doing the math in 16 bits. For FP8, vLLM’s docs say “Turing/Ampere GPUs are supported for W8A16 (weight-only FP8) utilizing Marlin kernels”. The memory savings stay, and the prefill speedup goes away.

The FP4 checkpoint on the H100 shows this. vLLM loaded it and logged that the “GPU does not have native support for FP4 computation but FP4 quantization is being used. Weight-only FP4 compression will be used leveraging the Marlin kernel.” The result, from the tables above:

FP4 on an H100, against BF16 on the same H100
weight memory              6.4 GB vs 16.4 GB     smaller, as on the B200
ms per token, 1 user       3.99 vs 6.60          1.65x faster: fewer bytes to read
ms per token, 16 users     about 18.0 vs 17.1    about 5% slower
first token, 8k prompt     about 437 vs 250 ms   1.75x slower

With 1 user, reading fewer bytes wins. With more work per step, converting every FP4 weight back to 16 bits on every step costs more than it saves. So an FP4 checkpoint is a good fit for a B200, and a poor one for an H100 serving long prompts. The same checkpoint can be either.

Accuracy: the cost grows as bits shrink Link to heading

Fewer values means every weight is rounded further from its trained value. Measured on the B200: all 1,319 GSM8K test questions (grade-school math word problems), thinking off, greedy decoding, one pass per format.

                  BF16              FP8               FP4
GSM8K correct     1,233  (93.5%)    1,231  (93.3%)    1,205  (91.4%)

The BF16 weights with an FP8 KV cache got 1,235. Part 2 ran BF16 and FP4 once before, on another B200, and got 1,231 and 1,205. So BF16 moved by 2 questions between runs. FP8 and the FP8 cache are within that noise. FP4 was about 26 lower both times. Harder tests lose more. RedHat’s evals of the FP4 checkpoint kept 99.4% of the BF16 score on GSM8K, but 79% on MMLU-Pro. Any format change needs an eval on your own workload.

Reading checkpoint names and flags Link to heading

  • No suffix, like Qwen/Qwen3-8B, usually means BF16. Check torch_dtype in the model’s config.json.
  • -FP8, -NVFP4, -MXFP4 mean the weights are stored in that format with scales. The details are in quantization_config in config.json. vLLM reads it from there.
  • --quantization usually stays unset. vLLM’s docs say that then “we first check the quantization_config attribute in the model config file”.
  • --dtype sets the format for weights and activations that aren’t quantized. The default, auto, uses “BF16 precision for BF16 models”.
  • --kv-cache-dtype sets the KV cache format, separately from the weights. auto matches the model. fp8 halves the cache’s memory.
  • Some models ship small from the start. DeepSeek-V3 was trained in FP8, and its card says “we only provide FP8 weights”. OpenAI’s gpt-oss stores its MoE weights in MXFP4.

Caveats Link to heading

  • Every measured number here is one run of one model, Qwen3-8B, on one B200 and one H100 with vLLM 0.30. Each speed test ran twice, and the tables show both runs. A larger model, a different batch size, other kernels or another vLLM version would change the numbers. The arithmetic doesn’t change.
  • The speed tests used random prompts with forced 512-token answers, not real traffic.
  • GSM8K ran once per format.

Next Link to heading

Part 2 takes FP4 apart: how 16 values can hold a model, where the missing speed goes, and what it costs.