vLLM, SGLang, TensorRT-LLM, Dynamo, llm-d, Fireworks. These names get compared as if they were the same kind of thing, but they sit at different layers. This post maps the layers, opens up the one layer they’re named after, the engine, and ends with what to check before running one in production. Part 2 covers deploying one, and part 3 the proprietary stacks you can’t see.

The layers Link to heading

Your app: sends /v1/chat/completions Hosted API Fireworks, Together, ... Runs every layer to the right for you. You see the API, not the engine, kernels or GPUs. Orchestration Dynamo, llm-d, Ray Serve: routing, autoscaling, multi-node Engine vLLM, SGLang, TensorRT-LLM: batching, KV cache, scheduling Kernels FlashAttention, FlashInfer, vendor kernels: GPU code Hardware NVIDIA, AMD, TPU, ... (or custom chips: Groq, Cerebras)
The layers of a serving stack. Run it yourself and you pick each layer. Use a hosted API and the provider picks all of them.

A replica is one running copy of a model, on as many GPUs as that copy needs, which mostly comes down to memory. The engine runs one replica: it takes requests, batches them, and runs them on the GPUs. A service runs several replicas behind a router when it needs more traffic than one copy can serve, or has to survive a failure. Part 2 shows how to pick the number.

Orchestration sits above the engine and manages those replicas. Dynamo’s README draws the same line: “it doesn’t replace SGLang, TensorRT-LLM, or vLLM, it turns them into a coordinated multi-node inference system”. It also says “If you’re running a single model on a single GPU, your inference engine alone is probably sufficient.”

Inference in production covered the orchestration layer: routing and autoscaling, failures and rollouts. This series is about the box below it.

Inside an engine Link to heading

API server OpenAI-compatible, tokenizes text Scheduler picks which requests run in the next step KV cache manager holds each request's KV cache in GPU memory, reuses it when a new prompt starts the same way Model runner runs every layer once for all requests in the batch GPU weights + KV cache in memory one new token per request per step, streamed back reads and writes the cache
batch of requests tokens cache bookkeeping
What every engine has inside. The loop runs once per step: the scheduler picks a batch, the runner does one forward pass, each request gets one token.

Every engine has these four parts. They differ in how each part works.

  • API server. Both vLLM and SGLang expose OpenAI-compatible endpoints like /v1/chat/completions, so a client written for one works with the other.
  • Scheduler. It decides which requests join each step. Continuous batching lets a new request join between steps. Chunked prefill splits a long prompt across steps, which Inference in production part 5 measured.
  • KV cache manager. It hands out GPU memory in blocks, and keeps prompts’ caches around for reuse. vLLM’s version is PagedAttention, and SGLang’s is RadixAttention.
  • Model runner. It runs one forward pass for the whole batch. Kernels, CUDA graphs and quantized weights live here, along with the formats from Weight precision part 1.

vLLM and SGLang Link to heading

vLLM and SGLang, side by side:

vLLM SGLang
Started at UC Berkeley, 2023 (PagedAttention paper) LMSYS, 2023 (RadixAttention paper)
Hosted by PyTorch Foundation LMSYS, a non-profit
Known for breadth: “200+ model architectures”, many hardware backends prefix caching, a “zero-overhead CPU scheduler”
Hardware NVIDIA, AMD, Intel, CPUs, plus plugins (TPU, Gaudi, …) NVIDIA, AMD, TPU, Intel, Apple, Ascend, …
API OpenAI-compatible OpenAI-compatible

Sources: the vLLM and SGLang READMEs.

They share more than the table suggests. SGLang’s README says “We learned the design and reused code from the following projects: Guidance, vLLM, LightLLM, FlashInfer, Outlines, and LMQL.” Both list the same big features: continuous batching, chunked prefill, prefix caching, speculative decoding, prefill/decode disaggregation, and FP8 and FP4 formats.

Three more names come up:

  • TensorRT-LLM is NVIDIA’s engine, for NVIDIA GPUs only.
  • TGI from Hugging Face is in maintenance mode since December 2025. Its README now recommends vLLM and SGLang.
  • llama.cpp and Ollama target laptops and small machines, not datacenter serving.

Which one is faster? Link to heading

Each project publishes benchmarks where it wins, and each one describes the versions it ran at the time.

SemiAnalysis, which benchmarks both every night, chose not to pick: “To prevent a restart of the SGLang vs vLLM benchmark wars and to save compute time, we decided to first pick only one of vLLM or SGLang as the default engine for each model.”

The answer depends on the model, the GPU, the traffic and the release. Test your own.

Before you pick one for production Link to heading

Check Why it matters Where to look
Your model runs on day one New architectures land in one engine before the other supported-models list, release notes
Your GPUs and formats An FP4 checkpoint needs FP4 support in the engine and the GPU hardware and quantization docs
Your traffic shape Shared prefixes reward prefix caching. Long prompts need chunked prefill or a split. your own request logs
More than one node Disaggregation and expert parallelism differ by engine. Or add an orchestration layer. engine docs, Dynamo, llm-d
Metrics and API You’ll need queue depth and KV use for autoscaling, and a stable API /metrics endpoint, API docs
Project health TGI shows an engine can stop moving release cadence, who hosts it

Pin the engine version, and re-test when you upgrade. Defaults change between releases, and so do benchmark results.

Next Link to heading

Part 2 deploys an engine: the container, the pod, splitting a model across GPUs, and what to watch once it runs.