Part 1 opened up an engine. This part runs one in production: what the pod looks like, what happens when the model doesn’t fit one GPU, and what to watch once it runs. The examples use vLLM on Kubernetes, with SGLang’s differences in a table. Part 3 covers stacks you can’t deploy yourself.

One replica Link to heading

Service port 80 → 8000 GPU node pod Engine container image: vllm/vllm-openai :8000 /v1/chat/completions :8000 /health → probes :8000 /metrics → Prometheus limits: nvidia.com/gpu: 1 Model cache PVC, weights /dev/shm in-memory volume GPU (weights + KV cache in its memory) Prometheus scrapes /metrics Hugging Face weights, first boot
One replica: a pod with one engine container on a GPU node. The model cache keeps weights across restarts. /dev/shm is shared memory the engine needs for tensor parallelism.

An engine ships as a container image: vllm/vllm-openai for vLLM, whose entrypoint is vllm serve. One replica is one pod running it. This is vLLM’s own Kubernetes example, cut down and pointed at Qwen3-8B:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: qwen3-8b
spec:
  replicas: 1
  selector:
    matchLabels: {app: qwen3-8b}
  template:
    metadata:
      labels: {app: qwen3-8b}
    spec:
      volumes:
      - name: cache-volume
        persistentVolumeClaim: {claimName: qwen3-8b}   # a 50Gi PVC in the docs
      - name: shm
        emptyDir: {medium: Memory, sizeLimit: "2Gi"}
      containers:
      - name: qwen3-8b
        image: vllm/vllm-openai:latest                  # pin a version in production
        command: ["/bin/sh", "-c"]
        args: ["vllm serve Qwen/Qwen3-8B"]
        ports:
        - containerPort: 8000
        resources:
          limits: {cpu: "10", memory: 20G, nvidia.com/gpu: "1"}
          requests: {cpu: "2", memory: 6G, nvidia.com/gpu: "1"}
        volumeMounts:
        - {name: cache-volume, mountPath: /root/.cache/huggingface}
        - {name: shm, mountPath: /dev/shm}
        livenessProbe:
          httpGet: {path: /health, port: 8000}
          initialDelaySeconds: 60
          periodSeconds: 10
        readinessProbe:
          httpGet: {path: /health, port: 8000}
          initialDelaySeconds: 60
          periodSeconds: 5

A Service in front maps port 80 to 8000. What each piece is for:

  • nvidia.com/gpu: "1" asks the scheduler for one GPU. The node needs NVIDIA’s device plugin.
  • The model cache keeps downloaded weights across pod restarts. The docs call the PVC optional, “you can use hostPath or other storage options”. Without it, every new pod downloads the model again.
  • /dev/shm is an in-memory volume. The manifest’s own comment says “vLLM needs to access the host’s shared memory for tensor parallel inference.”
  • The probes both hit /health. A first boot downloads about 16 GB and compiles, which took minutes in my measurements. The docs warn that a probe failing too early makes “Kubernetes scheduler kill the container”, and suggest raising failureThreshold.
  • The resource limits are the docs’ values, written for a 7B model. Check the memory limit for yours.

Bigger than one GPU Link to heading

Fits one GPU 1 pod whole model no flags more users: more replicas Fits one node 1 pod, 4 GPUs, NVLink 1/4 1/4 1/4 1/4 every layer split 4 ways --tensor-parallel-size 4 GPUs talk on every layer Needs two nodes 1 LeaderWorkerSet group (or a Ray cluster) leader pod, node 1 layers 1-18 TP across its GPUs worker pod, node 2 layers 19-36 TP across its GPUs TP = GPUs per node, PP = number of nodes the group starts, restarts and scales as one
Three ways to place one model. Tensor parallelism (TP) splits every layer across GPUs, which needs fast links, so it stays inside a node. Pipeline parallelism (PP) gives each node a slice of the layers.

How many GPUs one replica needs is a memory question. The weights have to fit, plus enough KV cache for the requests you want in flight, plus the engine’s own working memory:

GPUs per replica ≥ (weights + KV cache you want + engine overhead) / memory per GPU

Qwen3-8B in BF16 on an 80 GB H100:  16.4 GB of weights → 1 GPU
                                    vLLM gave 57 GB to the cache: 388,560 tokens,
                                    about 44 requests at 8k context
A 70B model in BF16:                140 GB of weights → 2 H100s just to hold them,
                                    more for the cache
The same 70B in FP8:                 70 GB → fits on 1 H100, with about 10 GB for the cache

The numbers for Qwen3-8B come from Weight precision part 1, which also covers the other lever: a smaller weight format halves the weights before you add a GPU. Memory sets the minimum. You can use more GPUs than that, because tensor parallelism also spreads each step’s weight reads across more memory bandwidth. I haven’t measured how much that speeds up each token.

Once you know the count, vLLM’s parallelism guide gives the rule for splitting in three lines:

The model fits Do this vLLM flags
on one GPU nothing: “distributed inference is probably unnecessary” none
on one node tensor parallelism --tensor-parallel-size = GPUs
only across nodes both TP = “GPUs per node”, PP = “number of nodes”

One exception: on GPUs without NVLink, “leverage pipeline parallelism instead of tensor parallelism”. Tensor parallelism exchanges data on every layer, so it needs the fastest link, which is why it stays inside one node.

Within one node, the engine starts one process per GPU itself. Across nodes, the pods have to start, fail and restart together. Two ways to get that on Kubernetes:

  • A Ray cluster, vLLM’s default for multi-node, which KubeRay can run.
  • LeaderWorkerSet, a Kubernetes API for “deploying a group of pods as a unit of replication”. Each group is “a single unit for rolling update, scaling”, with “all-or-nothing restart for failure handling”. vLLM, SGLang and KServe all document deployments with it.

How many replicas Link to heading

One replica is enough for development, or for traffic that fits in one replica when an outage during a restart is acceptable. Production needs more for any of three reasons:

  • Traffic. One replica serves only so many requests at once before every user’s tokens get too slow.
  • Failures. With one replica, any GPU or node failure is a full outage. Inference in production part 3 covers how often that happens.
  • Rollouts. A rolling update takes replicas out of service while new ones start, which takes minutes per replica.

To pick the number, measure what one replica can carry, then divide:

replicas = peak concurrent requests / requests one replica carries within your latency target
         + spare replicas for one failure and for a rollout

What one replica carries has two limits. Memory caps how many requests fit at all. From Weight precision part 1, vLLM had room for this much KV cache with Qwen3-8B in BF16:

B200: 1,056,992 tokens / (8,192 + 512 tokens per request) = about 121 requests at 8k context
H100:   388,560 tokens / (8,192 + 512 tokens per request) = about 44 requests

Latency can cap it sooner. More requests in a batch make every user’s tokens slower, as Inference in production part 5 showed. In one run with 16 requests at 8k context, each user got a token every 8.2 ms on the B200 and every 17.1 ms on the H100. I didn’t measure higher loads. Finding the real limit for your target takes a sweep: run vllm bench serve with your own prompt lengths at rising concurrency, and stop where p99 latency crosses your target.

A worked example, if 16 requests per replica is the limit: 100 concurrent requests at peak need 100 / 16 = 6.25, so 7 replicas. Add one for a failure and one for a rolling update, and that’s 9. The autoscaler can trim that off-peak, but a new replica takes minutes to arrive, so the spares have to exist before the peak, not after it.

What to watch Link to heading

Both engines expose a health endpoint and Prometheus metrics. The ones that matter first:

Need vLLM SGLang
Is it up? /health /health, or /health_generate, which generates one token
Requests queued vllm:num_requests_waiting sglang:num_queue_reqs
Requests running vllm:num_requests_running sglang:num_running_reqs
KV cache full? vllm:kv_cache_usage_perc sglang:token_usage
Time to first token vllm:time_to_first_token_seconds sglang:time_to_first_token_seconds

vLLM serves /metrics by default. SGLang needs --enable-metrics.

Two lessons from earlier posts apply here. A frozen replica can leave a plain health check hanging instead of failing. With default probe settings, a hung pod leaves the pool about 30 seconds after it stops answering. SGLang’s /health_generate checks the GPU by using it. And the queue and KV numbers are what autoscaling should follow, not CPU.

Shutting down Link to heading

vLLM’s --shutdown-timeout defaults to 0, which aborts requests in flight on SIGTERM. Part 4 of Inference in production measured what that does to a rollout, and the fix: set the timeout to about the p99 request time, and terminationGracePeriodSeconds above it.

SGLang’s differences Link to heading

vLLM SGLang
Image vllm/vllm-openai lmsysorg/sglang
Start vllm serve MODEL sglang serve MODEL --host 0.0.0.0
Port 8000 30000
Listens on all interfaces 127.0.0.1 unless you pass --host 0.0.0.0
Tensor parallel --tensor-parallel-size --tp-size
Metrics on --enable-metrics

SGLang’s server arguments list the defaults. Its repo has a single-node manifest close to the one above.

Beyond one deployment Link to heading

A Deployment and a Service get one model running. The orchestration layer from part 1 packages the rest:

Tool What it adds
production-stack a Helm chart: vLLM pods, a router for “KV cache reuse”, and Prometheus + Grafana
KServe a Hugging Face runtime that “uses vLLM backend engine”, and a newer CRD “built on the foundation of llm-d”
llm-d Helm recipes for a router, an InferencePool (“an LLM-optimized Service”) and vLLM or SGLang
Dynamo an operator for vLLM, SGLang or TensorRT-LLM, split prefill and decode, multi-node

Inference in production part 2 covered what the router and autoscaler in these do.

Weights at scale Link to heading

Every new replica needs the weights before it can serve. Dynamo’s install guide puts the cost plainly: without shared storage, “large models (>70B) take hours to download per pod, and many replicas will hit HuggingFace rate limits.” The usual answers:

  • A shared volume that all pods mount, so the download happens once.
  • Streaming from object storage. vLLM’s --load-format runai_streamer reads s3://, gs:// or az:// paths straight into GPU memory.
  • Keeping the compile cache. vLLM stores its compiled kernels under VLLM_CACHE_ROOT, and its docs say they “can be copied between machines or baked into a container image”. In my cold-start measurements, keeping them saved over two minutes on a 7B model.

Next Link to heading

Part 3 looks at the stacks you can’t deploy yourself: managed open engines, proprietary engines like Fireworks’, and custom chips, and how to evaluate them.