Part 1 opened up an engine. This part runs one in production: what the pod looks like, what happens when the model doesn’t fit one GPU, and what to watch once it runs. The examples use vLLM on Kubernetes, with SGLang’s differences in a table. Part 3 covers stacks you can’t deploy yourself.
One replica Link to heading
An engine ships as a container image: vllm/vllm-openai for vLLM, whose entrypoint is vllm serve. One replica is one pod running it. This is vLLM’s own Kubernetes example, cut down and pointed at Qwen3-8B:
apiVersion: apps/v1
kind: Deployment
metadata:
name: qwen3-8b
spec:
replicas: 1
selector:
matchLabels: {app: qwen3-8b}
template:
metadata:
labels: {app: qwen3-8b}
spec:
volumes:
- name: cache-volume
persistentVolumeClaim: {claimName: qwen3-8b} # a 50Gi PVC in the docs
- name: shm
emptyDir: {medium: Memory, sizeLimit: "2Gi"}
containers:
- name: qwen3-8b
image: vllm/vllm-openai:latest # pin a version in production
command: ["/bin/sh", "-c"]
args: ["vllm serve Qwen/Qwen3-8B"]
ports:
- containerPort: 8000
resources:
limits: {cpu: "10", memory: 20G, nvidia.com/gpu: "1"}
requests: {cpu: "2", memory: 6G, nvidia.com/gpu: "1"}
volumeMounts:
- {name: cache-volume, mountPath: /root/.cache/huggingface}
- {name: shm, mountPath: /dev/shm}
livenessProbe:
httpGet: {path: /health, port: 8000}
initialDelaySeconds: 60
periodSeconds: 10
readinessProbe:
httpGet: {path: /health, port: 8000}
initialDelaySeconds: 60
periodSeconds: 5
A Service in front maps port 80 to 8000. What each piece is for:
nvidia.com/gpu: "1"asks the scheduler for one GPU. The node needs NVIDIA’s device plugin.- The model cache keeps downloaded weights across pod restarts. The docs call the PVC optional, “you can use hostPath or other storage options”. Without it, every new pod downloads the model again.
/dev/shmis an in-memory volume. The manifest’s own comment says “vLLM needs to access the host’s shared memory for tensor parallel inference.”- The probes both hit
/health. A first boot downloads about 16 GB and compiles, which took minutes in my measurements. The docs warn that a probe failing too early makes “Kubernetes scheduler kill the container”, and suggest raisingfailureThreshold. - The resource limits are the docs’ values, written for a 7B model. Check the memory limit for yours.
Bigger than one GPU Link to heading
How many GPUs one replica needs is a memory question. The weights have to fit, plus enough KV cache for the requests you want in flight, plus the engine’s own working memory:
GPUs per replica ≥ (weights + KV cache you want + engine overhead) / memory per GPU
Qwen3-8B in BF16 on an 80 GB H100: 16.4 GB of weights → 1 GPU
vLLM gave 57 GB to the cache: 388,560 tokens,
about 44 requests at 8k context
A 70B model in BF16: 140 GB of weights → 2 H100s just to hold them,
more for the cache
The same 70B in FP8: 70 GB → fits on 1 H100, with about 10 GB for the cache
The numbers for Qwen3-8B come from Weight precision part 1, which also covers the other lever: a smaller weight format halves the weights before you add a GPU. Memory sets the minimum. You can use more GPUs than that, because tensor parallelism also spreads each step’s weight reads across more memory bandwidth. I haven’t measured how much that speeds up each token.
Once you know the count, vLLM’s parallelism guide gives the rule for splitting in three lines:
| The model fits | Do this | vLLM flags |
|---|---|---|
| on one GPU | nothing: “distributed inference is probably unnecessary” | none |
| on one node | tensor parallelism | --tensor-parallel-size = GPUs |
| only across nodes | both | TP = “GPUs per node”, PP = “number of nodes” |
One exception: on GPUs without NVLink, “leverage pipeline parallelism instead of tensor parallelism”. Tensor parallelism exchanges data on every layer, so it needs the fastest link, which is why it stays inside one node.
Within one node, the engine starts one process per GPU itself. Across nodes, the pods have to start, fail and restart together. Two ways to get that on Kubernetes:
- A Ray cluster, vLLM’s default for multi-node, which KubeRay can run.
- LeaderWorkerSet, a Kubernetes API for “deploying a group of pods as a unit of replication”. Each group is “a single unit for rolling update, scaling”, with “all-or-nothing restart for failure handling”. vLLM, SGLang and KServe all document deployments with it.
How many replicas Link to heading
One replica is enough for development, or for traffic that fits in one replica when an outage during a restart is acceptable. Production needs more for any of three reasons:
- Traffic. One replica serves only so many requests at once before every user’s tokens get too slow.
- Failures. With one replica, any GPU or node failure is a full outage. Inference in production part 3 covers how often that happens.
- Rollouts. A rolling update takes replicas out of service while new ones start, which takes minutes per replica.
To pick the number, measure what one replica can carry, then divide:
replicas = peak concurrent requests / requests one replica carries within your latency target
+ spare replicas for one failure and for a rollout
What one replica carries has two limits. Memory caps how many requests fit at all. From Weight precision part 1, vLLM had room for this much KV cache with Qwen3-8B in BF16:
B200: 1,056,992 tokens / (8,192 + 512 tokens per request) = about 121 requests at 8k context
H100: 388,560 tokens / (8,192 + 512 tokens per request) = about 44 requests
Latency can cap it sooner. More requests in a batch make every user’s tokens slower, as Inference in production part 5 showed. In one run with 16 requests at 8k context, each user got a token every 8.2 ms on the B200 and every 17.1 ms on the H100. I didn’t measure higher loads. Finding the real limit for your target takes a sweep: run vllm bench serve with your own prompt lengths at rising concurrency, and stop where p99 latency crosses your target.
A worked example, if 16 requests per replica is the limit: 100 concurrent requests at peak need 100 / 16 = 6.25, so 7 replicas. Add one for a failure and one for a rolling update, and that’s 9. The autoscaler can trim that off-peak, but a new replica takes minutes to arrive, so the spares have to exist before the peak, not after it.
What to watch Link to heading
Both engines expose a health endpoint and Prometheus metrics. The ones that matter first:
| Need | vLLM | SGLang |
|---|---|---|
| Is it up? | /health |
/health, or /health_generate, which generates one token |
| Requests queued | vllm:num_requests_waiting |
sglang:num_queue_reqs |
| Requests running | vllm:num_requests_running |
sglang:num_running_reqs |
| KV cache full? | vllm:kv_cache_usage_perc |
sglang:token_usage |
| Time to first token | vllm:time_to_first_token_seconds |
sglang:time_to_first_token_seconds |
vLLM serves /metrics by default. SGLang needs --enable-metrics.
Two lessons from earlier posts apply here. A frozen replica can leave a plain health check hanging instead of failing. With default probe settings, a hung pod leaves the pool about 30 seconds after it stops answering. SGLang’s /health_generate checks the GPU by using it. And the queue and KV numbers are what autoscaling should follow, not CPU.
Shutting down Link to heading
vLLM’s --shutdown-timeout defaults to 0, which aborts requests in flight on SIGTERM. Part 4 of Inference in production measured what that does to a rollout, and the fix: set the timeout to about the p99 request time, and terminationGracePeriodSeconds above it.
SGLang’s differences Link to heading
| vLLM | SGLang | |
|---|---|---|
| Image | vllm/vllm-openai |
lmsysorg/sglang |
| Start | vllm serve MODEL |
sglang serve MODEL --host 0.0.0.0 |
| Port | 8000 | 30000 |
| Listens on | all interfaces | 127.0.0.1 unless you pass --host 0.0.0.0 |
| Tensor parallel | --tensor-parallel-size |
--tp-size |
| Metrics | on | --enable-metrics |
SGLang’s server arguments list the defaults. Its repo has a single-node manifest close to the one above.
Beyond one deployment Link to heading
A Deployment and a Service get one model running. The orchestration layer from part 1 packages the rest:
| Tool | What it adds |
|---|---|
| production-stack | a Helm chart: vLLM pods, a router for “KV cache reuse”, and Prometheus + Grafana |
| KServe | a Hugging Face runtime that “uses vLLM backend engine”, and a newer CRD “built on the foundation of llm-d” |
| llm-d | Helm recipes for a router, an InferencePool (“an LLM-optimized Service”) and vLLM or SGLang |
| Dynamo | an operator for vLLM, SGLang or TensorRT-LLM, split prefill and decode, multi-node |
Inference in production part 2 covered what the router and autoscaler in these do.
Weights at scale Link to heading
Every new replica needs the weights before it can serve. Dynamo’s install guide puts the cost plainly: without shared storage, “large models (>70B) take hours to download per pod, and many replicas will hit HuggingFace rate limits.” The usual answers:
- A shared volume that all pods mount, so the download happens once.
- Streaming from object storage. vLLM’s
--load-format runai_streamerreadss3://,gs://oraz://paths straight into GPU memory. - Keeping the compile cache. vLLM stores its compiled kernels under
VLLM_CACHE_ROOT, and its docs say they “can be copied between machines or baked into a container image”. In my cold-start measurements, keeping them saved over two minutes on a 7B model.
Next Link to heading
Part 3 looks at the stacks you can’t deploy yourself: managed open engines, proprietary engines like Fireworks’, and custom chips, and how to evaluate them.