An LLM running on a GPU spends its time in kernels, the small programs the GPU runs one after another. In the measurement later in this post, generating one token with Qwen3-8B on an H100 took 2,398 of them. This post starts from the beginning: what a kernel is, how model code turns into kernels, where kernels come from, and which ones run for each token.
Why a GPU Link to heading
Generating one token means multiplying a vector by every weight matrix in the model. Each output number is a long sum of multiplications, and every output number can be worked out at the same time as the others. A CPU has a handful of powerful cores and does a few of these at a time. A GPU has thousands of simpler ones and does many at once.
The two numbers that matter for an H100, from Sizing inference: it can do 989 TFLOPS (trillion operations per second) of BF16 math, and its memory moves 3.35 TB/s. With one user, generating a token mostly needs the second. Each step reads nearly every weight once (15.1 GB of Qwen3-8B’s 16.4 GB) and uses each one for a single multiply and add. In BF16 a weight is 2 bytes, so that’s 2 operations for every 2 bytes read. Qwen3-8B multiplies 7.57 billion weights per token:
math 2 × 7.57 billion = 15.1 billion operations / 989 TFLOPS = 0.015 ms
memory 15.1 GB of weights / 3.35 TB/s = 4.5 ms
The math takes about 0.3% of the time it takes to read the weights, so the speed of memory decides how fast a token can come out: 4.5 ms at best. When many requests share a step, each weight read is reused for all of them, and Sizing works out that math only becomes the limit at around 295 requests.
What a kernel is Link to heading
The GPU doesn’t run a whole program on its own. The CPU runs the program and hands the GPU one piece of work at a time. Each piece is a kernel: a function written to run on the GPU across many threads at once.
Adding two vectors of a million numbers is one kernel. The CPU launches it, telling the GPU which function to run, where the two vectors and the result live in GPU memory, and how many threads to use. Each thread adds one pair of numbers. A matrix multiply is another kernel, where each group of threads works out one block of the output.
The CPU doesn’t wait for a kernel to finish before launching the next one. It puts launches in a queue, the GPU works through the queue in order, and the CPU keeps adding to it.
From model code to kernels Link to heading
Model code doesn’t mention kernels. It’s written with operations in a framework like PyTorch, which Hugging Face transformers, used for the measurement later in this post, runs on:
gate = x @ W_gate # a matrix multiply
up = x @ W_up # another matrix multiply
h = silu(gate) * up # an activation, then an elementwise multiply
When this runs on a GPU, PyTorch picks a kernel for each operation and launches it. These three lines are at least four kernels: two matrix multiplies, one for silu, and one for the multiply. A library can split one operation into more than one kernel, as cuBLAS does with some matrix multiplies. In plain PyTorch, every operation is at least one kernel, so how a model is written decides how many kernels run. Other frameworks, such as JAX, and C++ projects such as llama.cpp, work the same way: operations become kernels. RMSNorm, the normalization in every layer of models like Qwen3, takes one line to describe and is written as several operations (square, take the mean, add a small number, take the reciprocal square root, multiply), so it runs as several kernels.
Where kernels come from Link to heading
Nobody writes most of these by hand for each model. Kernels come from a few places:
- Vendor libraries. NVIDIA’s cuBLAS has matrix multiply kernels and cuDNN has kernels for neural network layers, including attention. PyTorch calls them for those operations.
- PyTorch’s own kernels, for the long tail of simpler operations: adding, multiplying, copying, converting between number formats.
- Hand-written kernels for one job. FlashAttention is the best-known example: an attention kernel written to read and write GPU memory as little as possible.
- Generated kernels. Triton lets you write kernels in Python instead of CUDA C++.
torch.compilegenerates kernels from model code, and can combine several operations into one kernel, so the data is read once instead of once per operation. - The engine’s own. vLLM, SGLang and TensorRT-LLM ship kernels written for serving, and choose among all of the above for each operation.
The kernels in one token Link to heading
To see what this adds up to, I profiled one decode step of Qwen3-8B in BF16 on an H100, run with Hugging Face transformers, which runs the model’s PyTorch code as written. The setup is in the lab repo. Generating one token launched 2,398 kernels. Grouped by kind:
kind kernels GPU time share
matrix multiply 433 6.20 ms 61%
elementwise 1,234 1.97 ms 20%
copy or concat 547 1.28 ms 13%
reduction 146 0.36 ms 4%
attention 36 0.29 ms 3%
other 2 0.00 ms 0%
total 2,398 10.09 ms
- Matrix multiplies are the seven weight matrices in each of the 36 layers, plus the output layer that scores every possible next token, and some helper kernels cuBLAS adds. They’re the fewest kernels with the most time, because they read the weights.
- Elementwise kernels do one small operation on every number of a tensor: an activation, an add, a multiply.
- Copy or concat kernels move data or convert it between number formats, for example from BF16 to 32-bit floats for a norm and back.
- Reductions collapse many numbers into one, like the mean in RMSNorm or picking the most likely token.
- Attention is one kernel per layer, from cuDNN here. At this short context (a 256-token prompt) it’s 3% of the time. With longer contexts it reads a bigger KV cache and takes more.
This also shows how good the matrix multiply kernels are. They’re the kernels that read the 15.1 GB of weights, so the 4.5 ms worked out earlier is the fastest they could possibly go. They took 6.2 ms. Dividing the bytes by the time they took gives the speed they actually read memory at:
best case 15.1 GB at 3.35 TB/s = 4.5 ms
measured 15.1 GB in 6.2 ms = 2.4 TB/s, 73% of 3.35
So the matrix multiplies spent 6.2 ms on work the memory could feed in 4.5 ms. That 1.7 ms gap is room for a better kernel.
Why inference cares Link to heading
How fast tokens come out depends on three things about these kernels:
-
How close each kernel gets to the hardware’s limit. The matrix multiplies here reached 73% of memory bandwidth. Better kernels read the same bytes faster.
-
How many kernels there are. Almost 2,000 of the 2,398 are small elementwise, copy and reduction kernels, about 3.9 ms of GPU time together. Each one reads its input from memory and writes its output back, so combining several into one kernel saves both the trips and the launches.
-
How they’re launched. In this run, the GPU sat idle for about two thirds of every token, waiting for the CPU to tell it what to run next. The GPU doesn’t pick its next kernel itself. The CPU launches each one, and here each of the 2,398 launches went through Python and PyTorch, which took longer than the GPU needed to run the kernel:
token took 29.3 ms GPU working 10.1 ms GPU waiting 19.3 ms a new kernel arrived every 29.3 ms / 2,398 = 12.2 µs GPU finished each in 10.1 ms / 2,398 = 4.2 µs on averageA faster GPU wouldn’t help here; the launches have to arrive faster. CUDA graphs are the usual fix: they record the whole sequence of launches once, then replay it with a single launch.
Caveats Link to heading
- One setup, two runs. One H100, batch 1, a 256-token prompt, transformers 5.19 and PyTorch 2.14. The kernel counts and GPU time matched between the runs, but the step took 29.3 ms in one and 37.4 ms in the other, because it depended on how fast the CPU launched. A serving engine like vLLM runs the same model with fewer, larger kernels.
- The kinds come from kernel names, matched by pattern, so the grouping is approximate.