A hosted model answers thousands of people at once, and a single user still gets their tokens quickly. Both come from batching, where one GPU works on many requests in the same step. This post explains why batching is nearly free at first, measures what happens as one H100 takes on more users, and turns the result into the cost of a token.

Why batching is nearly free at first Link to heading

Generating a token reads the model’s weights from GPU memory. Kernels worked out what that costs for Qwen3-8B on an H100: each step reads 15.1 GB of weights, and the memory moves 3.35 TB/s, so a step takes at least 4.5 ms. The math in that step takes a small fraction of the time; the GPU mostly waits for the weights to arrive.

With two users, the step still reads each weight once, but uses it twice: once for each user’s next token. With eight users, eight times. The weights are the expensive part, and they don’t get any bigger when more users share them. So the first users added to a step cost almost nothing, and each step produces more tokens.

That holds until the math catches up. Each extra user adds work to every weight that’s read, and at some point the GPU spends longer computing than waiting for memory. After that, more users make every step slower, and total throughput stops growing.

Engines don’t wait for a full batch to start, either. Inference engines part 1 described the scheduler: “Continuous batching lets a new request join between steps.” When one user’s answer finishes, the next request takes its place in the next step.

Setup Link to heading

  • Qwen3-8B in BF16 on one H100 80GB on Modal, served by vLLM 0.30, with prefix caching off.
  • Every request has a 512-token prompt and a 256-token answer, using vLLM’s bench serve with random tokens.
  • The number of users at once goes from 1 to 512, doubling each time. Each level sends at least 32 requests, and 4 per user above that.
  • vLLM had room for 390,816 tokens of KV cache, and a request here needs 768, so about 508 requests fit at once. No request was pushed out for lack of memory at any level.
  • Everything ran twice. The script and raw results are in the lab repo.

Results Link to heading

From run 1. In run 2, time per token and total output were within a few percent of these except at 512 users, noted below. Time to first token varied more between runs: 939 ms against 722 at 128 users, for example.

users   time per token   tokens/s per user   tokens/s total   time to first token
    1         6.85 ms                 146              145                 23 ms
    2         6.88 ms                 145              285                 41 ms
    4         7.01 ms                 143              556                 54 ms
    8         7.33 ms                 136            1,046                 86 ms
   16         8.11 ms                 123            1,820                179 ms
   32         8.53 ms                 117            3,306                299 ms
   64        10.57 ms                  95            5,136                488 ms
  128        15.16 ms                  66            7,121                722 ms
  256        23.64 ms                  42            8,458              1,690 ms
  512        47.69 ms                  21            8,555              3,071 ms

Time per token is the average gap between one user’s tokens. Tokens/s total is everything the GPU produced.

Reading down the table:

1 to 8 users     total  1,046 / 145   =  7.2x       each user  7.33 / 6.85 ms  =  7% slower
1 to 64 users    total  5,136 / 145   = 35.5x       each user 10.57 / 6.85 ms  = 54% slower
128 to 256       total  8,458 / 7,121 = +19%
256 to 512       total  8,555 / 8,458 =  +1%        (run 2: +7%)

Adding the first eight users gave seven times the output, with each user’s tokens 7% slower. By 64 users the GPU makes 35 times as many tokens, and each user still gets 95 tokens per second. Past 256 users, total output barely grows, and every added user only makes the others slower: at 512 users each one gets 21 tokens per second.

Why it stops growing Link to heading

At 512 users the GPU is no longer waiting on memory. Each request also has a 512-token prompt to prefill, so for every token it outputs, the GPU processes three: one new token and two prompt tokens. Each token takes about 2 operations per weight, and Qwen3-8B multiplies 7.57 billion weights (Sizing):

8,555 output tokens/s  x 3 tokens  x 2 x 7.57 billion  =  389 trillion operations per second
                                                       =  39% of the H100's 989 TFLOPS

That leaves out attention, so the real share is higher. Sizing found that “Real kernels reached about half of the datasheet’s peak.” At 512 users this GPU is close to that, and more users can’t make it produce more.

Time to first token rises too, from 23 ms to 3 seconds. Each new request waits for room in the batch and shares the GPU with everyone else’s prompts. Run 2’s 512-user first token came in at 1.9 seconds, so that number moves a lot between runs.

What a token costs Link to heading

Modal charges $0.001097 per second for an H100 as of October 2026, or $3.95 an hour, for the GPU alone. Dividing by the tokens it produced per second:

users   tokens/s total   cost per million output tokens
    1              145                            $7.59
    2              285                            $3.85
    4              556                            $1.97
    8            1,046                            $1.05
   16            1,820                            $0.60
   32            3,306                            $0.33
   64            5,136                            $0.21
  128            7,121                            $0.15
  256            8,458                            $0.13
  512            8,555                            $0.13

The same GPU costs the same per hour whether it serves one user or 512. With one user, a million tokens costs $7.59. With 64 users it costs 21 cents, and past 256 users it stays at about 13 cents.

A provider picks a point on this curve with a latency target. Sizing used one: “Each user should get at least 50 tokens per second, so at most 20 ms per token”. Here, 128 users get 15.16 ms per token and 256 get 23.64 ms, so the most this GPU can carry under 20 ms is somewhere between the two. At 128 users that’s 7,121 tokens per second, or 15 cents per million. Pushing to 256 users would save about 2 cents per million and break the target.

Caveats Link to heading

  • One setup, two runs. One H100, one model, 512 tokens in and 256 out, random tokens, and every answer forced to its full length. Longer prompts bring the math limit sooner, and longer contexts make the KV cache fill up sooner.
  • The GPU price is only the GPU. Modal also bills CPU and memory, and a real fleet carries idle time and spares, so these costs are a floor.
  • Run 2 differed at 512 users: 9,122 tokens per second against 8,555, and 1.9 seconds to first token against 3.1. At the other levels, time per token and total output agreed within a few percent, and time to first token less closely.