<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>SGLang on </title>
    <link>https://hiren.me/tags/sglang/</link>
    <description>Recent content in SGLang on </description>
    <generator>Hugo</generator>
    <language>en</language>
    <lastBuildDate>Mon, 05 Oct 2026 16:10:00 -0700</lastBuildDate>
    <atom:link href="https://hiren.me/tags/sglang/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Inference engines (part 2): deploying one</title>
      <link>https://hiren.me/posts/inference-engines-part-2/</link>
      <pubDate>Mon, 05 Oct 2026 16:10:00 -0700</pubDate>
      <guid>https://hiren.me/posts/inference-engines-part-2/</guid>
      <description>&lt;p&gt;&lt;a href=&#34;https://hiren.me/posts/inference-engines-part-1/&#34; &gt;Part 1&lt;/a&gt; opened up an engine. This part runs one in production: what the pod looks like, what happens when the model doesn&amp;rsquo;t fit one GPU, and what to watch once it runs. The examples use vLLM on Kubernetes, with SGLang&amp;rsquo;s differences in a table. &lt;a href=&#34;https://hiren.me/posts/inference-engines-part-3/&#34; &gt;Part 3&lt;/a&gt; covers stacks you can&amp;rsquo;t deploy yourself.&lt;/p&gt;&#xA;&lt;h2 id=&#34;one-replica&#34;&gt;&#xA;  One replica&#xA;  &lt;a class=&#34;heading-link&#34; href=&#34;#one-replica&#34;&gt;&#xA;    &lt;i class=&#34;fa-solid fa-link&#34; aria-hidden=&#34;true&#34; title=&#34;Link to heading&#34;&gt;&lt;/i&gt;&#xA;    &lt;span class=&#34;sr-only&#34;&gt;Link to heading&lt;/span&gt;&#xA;  &lt;/a&gt;&#xA;&lt;/h2&gt;&#xA;&lt;figure class=&#34;epod&#34;&gt;&#xA;  &lt;style&gt;&#xA;    .epod { --ep-req: #2a78d6; --ep-hi: #2a78d6; --ep-box: rgba(128,128,128,.10); --ep-edge: rgba(128,128,128,.55);&#xA;            border: 1px solid rgba(128,128,128,.35); border-radius: 8px; padding: 14px; margin: 1.5em 0; }&#xA;    @media (prefers-color-scheme: dark) { body.colorscheme-auto .epod { --ep-req: #3987e5; --ep-hi: #3987e5; } }&#xA;    body.colorscheme-dark .epod { --ep-req: #3987e5; --ep-hi: #3987e5; }&#xA;    .epod svg { width: 100%; height: auto; display: block; }&#xA;    .epod svg text { fill: currentColor; font-size: 12.5px; }&#xA;    .epod svg .t { font-weight: 700; font-size: 13px; }&#xA;    .epod svg .s { font-size: 10.5px; opacity: .8; }&#xA;    .epod svg rect { fill: var(--ep-box); stroke: var(--ep-edge); }&#xA;    .epod svg rect.hi { stroke: var(--ep-hi); stroke-width: 2; }&#xA;    .epod svg rect.grp { fill: none; stroke-dasharray: 5 4; }&#xA;    .epod svg .req { stroke: var(--ep-req); stroke-width: 2; fill: none; }&#xA;    .epod svg .ctl { stroke: var(--ep-edge); stroke-width: 1.5; stroke-dasharray: 5 4; fill: none; }&#xA;    .epod figcaption { font-size: 12.5px; opacity: .8; margin-top: 8px; }&#xA;  &lt;/style&gt;&#xA;  &lt;svg viewBox=&#34;0 0 760 300&#34; role=&#34;img&#34; aria-label=&#34;One engine replica on Kubernetes&#34;&gt;&#xA;    &lt;defs&gt;&#xA;      &lt;marker id=&#34;epod-r&#34; viewBox=&#34;0 0 10 10&#34; refX=&#34;9&#34; refY=&#34;5&#34; markerWidth=&#34;7&#34; markerHeight=&#34;7&#34; orient=&#34;auto&#34;&gt;&lt;path d=&#34;M0,0 L10,5 L0,10 z&#34; style=&#34;fill:var(--ep-req)&#34;/&gt;&lt;/marker&gt;&#xA;    &lt;/defs&gt;&#xA;    &lt;rect x=&#34;8&#34; y=&#34;110&#34; width=&#34;120&#34; height=&#34;64&#34; rx=&#34;6&#34;/&gt;&#xA;    &lt;text class=&#34;t&#34; x=&#34;68&#34; y=&#34;136&#34; text-anchor=&#34;middle&#34;&gt;Service&lt;/text&gt;&#xA;    &lt;text class=&#34;s&#34; x=&#34;68&#34; y=&#34;154&#34; text-anchor=&#34;middle&#34;&gt;port 80 → 8000&lt;/text&gt;&#xA;&#xA;    &lt;rect class=&#34;grp&#34; x=&#34;160&#34; y=&#34;20&#34; width=&#34;420&#34; height=&#34;270&#34; rx=&#34;8&#34;/&gt;&#xA;    &lt;text class=&#34;s&#34; x=&#34;172&#34; y=&#34;38&#34;&gt;GPU node&lt;/text&gt;&#xA;    &lt;rect class=&#34;grp&#34; x=&#34;176&#34; y=&#34;48&#34; width=&#34;390&#34; height=&#34;180&#34; rx=&#34;8&#34;/&gt;&#xA;    &lt;text class=&#34;s&#34; x=&#34;188&#34; y=&#34;66&#34;&gt;pod&lt;/text&gt;&#xA;&#xA;    &lt;rect class=&#34;hi&#34; x=&#34;196&#34; y=&#34;76&#34; width=&#34;210&#34; height=&#34;136&#34; rx=&#34;6&#34;/&gt;&#xA;    &lt;text class=&#34;t&#34; x=&#34;301&#34; y=&#34;100&#34; text-anchor=&#34;middle&#34;&gt;Engine container&lt;/text&gt;&#xA;    &lt;text class=&#34;s&#34; x=&#34;301&#34; y=&#34;118&#34; text-anchor=&#34;middle&#34;&gt;image: vllm/vllm-openai&lt;/text&gt;&#xA;    &lt;text class=&#34;s&#34; x=&#34;210&#34; y=&#34;146&#34;&gt;:8000 /v1/chat/completions&lt;/text&gt;&#xA;    &lt;text class=&#34;s&#34; x=&#34;210&#34; y=&#34;164&#34;&gt;:8000 /health → probes&lt;/text&gt;&#xA;    &lt;text class=&#34;s&#34; x=&#34;210&#34; y=&#34;182&#34;&gt;:8000 /metrics → Prometheus&lt;/text&gt;&#xA;    &lt;text class=&#34;s&#34; x=&#34;210&#34; y=&#34;200&#34;&gt;limits: nvidia.com/gpu: 1&lt;/text&gt;&#xA;&#xA;    &lt;rect x=&#34;430&#34; y=&#34;76&#34; width=&#34;122&#34; height=&#34;60&#34; rx=&#34;6&#34;/&gt;&#xA;    &lt;text class=&#34;t&#34; x=&#34;491&#34; y=&#34;100&#34; text-anchor=&#34;middle&#34;&gt;Model cache&lt;/text&gt;&#xA;    &lt;text class=&#34;s&#34; x=&#34;491&#34; y=&#34;118&#34; text-anchor=&#34;middle&#34;&gt;PVC, weights&lt;/text&gt;&#xA;&#xA;    &lt;rect x=&#34;430&#34; y=&#34;150&#34; width=&#34;122&#34; height=&#34;62&#34; rx=&#34;6&#34;/&gt;&#xA;    &lt;text class=&#34;t&#34; x=&#34;491&#34; y=&#34;174&#34; text-anchor=&#34;middle&#34;&gt;/dev/shm&lt;/text&gt;&#xA;    &lt;text class=&#34;s&#34; x=&#34;491&#34; y=&#34;192&#34; text-anchor=&#34;middle&#34;&gt;in-memory volume&lt;/text&gt;&#xA;&#xA;    &lt;rect x=&#34;196&#34; y=&#34;240&#34; width=&#34;370&#34; height=&#34;40&#34; rx=&#34;6&#34;/&gt;&#xA;    &lt;text class=&#34;t&#34; x=&#34;381&#34; y=&#34;265&#34; text-anchor=&#34;middle&#34;&gt;GPU (weights + KV cache in its memory)&lt;/text&gt;&#xA;&#xA;    &lt;rect x=&#34;610&#34; y=&#34;76&#34; width=&#34;142&#34; height=&#34;64&#34; rx=&#34;6&#34;/&gt;&#xA;    &lt;text class=&#34;t&#34; x=&#34;681&#34; y=&#34;102&#34; text-anchor=&#34;middle&#34;&gt;Prometheus&lt;/text&gt;&#xA;    &lt;text class=&#34;s&#34; x=&#34;681&#34; y=&#34;120&#34; text-anchor=&#34;middle&#34;&gt;scrapes /metrics&lt;/text&gt;&#xA;&#xA;    &lt;rect x=&#34;610&#34; y=&#34;160&#34; width=&#34;142&#34; height=&#34;64&#34; rx=&#34;6&#34;/&gt;&#xA;    &lt;text class=&#34;t&#34; x=&#34;681&#34; y=&#34;186&#34; text-anchor=&#34;middle&#34;&gt;Hugging Face&lt;/text&gt;&#xA;    &lt;text class=&#34;s&#34; x=&#34;681&#34; y=&#34;204&#34; text-anchor=&#34;middle&#34;&gt;weights, first boot&lt;/text&gt;&#xA;&#xA;    &lt;line class=&#34;req&#34; x1=&#34;128&#34; y1=&#34;142&#34; x2=&#34;194&#34; y2=&#34;142&#34; marker-end=&#34;url(#epod-r)&#34;/&gt;&#xA;    &lt;line class=&#34;ctl&#34; x1=&#34;406&#34; y1=&#34;106&#34; x2=&#34;428&#34; y2=&#34;106&#34;/&gt;&#xA;    &lt;line class=&#34;ctl&#34; x1=&#34;406&#34; y1=&#34;181&#34; x2=&#34;428&#34; y2=&#34;181&#34;/&gt;&#xA;    &lt;line class=&#34;ctl&#34; x1=&#34;301&#34; y1=&#34;212&#34; x2=&#34;301&#34; y2=&#34;238&#34;/&gt;&#xA;  &lt;/svg&gt;&#xA;  &lt;figcaption&gt;One replica: a pod with one engine container on a GPU node. The model cache keeps weights across restarts. /dev/shm is shared memory the engine needs for tensor parallelism.&lt;/figcaption&gt;&#xA;&lt;/figure&gt;&#xA;&#xA;&lt;p&gt;An engine ships as a container image: &lt;a href=&#34;https://docs.vllm.ai/en/latest/deployment/docker/&#34;  class=&#34;external-link&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;&lt;code&gt;vllm/vllm-openai&lt;/code&gt;&lt;/a&gt; for vLLM, whose entrypoint is &lt;code&gt;vllm serve&lt;/code&gt;. One replica is one pod running it. This is vLLM&amp;rsquo;s own &lt;a href=&#34;https://docs.vllm.ai/en/latest/deployment/k8s/&#34;  class=&#34;external-link&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Kubernetes example&lt;/a&gt;, cut down and pointed at Qwen3-8B:&lt;/p&gt;</description>
    </item>
    <item>
      <title>Inference engines (part 1): what an engine does and where it sits</title>
      <link>https://hiren.me/posts/inference-engines-part-1/</link>
      <pubDate>Mon, 05 Oct 2026 16:00:00 -0700</pubDate>
      <guid>https://hiren.me/posts/inference-engines-part-1/</guid>
      <description>&lt;p&gt;vLLM, SGLang, TensorRT-LLM, Dynamo, llm-d, Fireworks. These names get compared as if they were the same kind of thing, but they sit at different layers. This post maps the layers, opens up the one layer they&amp;rsquo;re named after, the engine, and ends with what to check before running one in production. &lt;a href=&#34;https://hiren.me/posts/inference-engines-part-2/&#34; &gt;Part 2&lt;/a&gt; covers deploying one, and &lt;a href=&#34;https://hiren.me/posts/inference-engines-part-3/&#34; &gt;part 3&lt;/a&gt; the proprietary stacks you can&amp;rsquo;t see.&lt;/p&gt;&#xA;&lt;h2 id=&#34;the-layers&#34;&gt;&#xA;  The layers&#xA;  &lt;a class=&#34;heading-link&#34; href=&#34;#the-layers&#34;&gt;&#xA;    &lt;i class=&#34;fa-solid fa-link&#34; aria-hidden=&#34;true&#34; title=&#34;Link to heading&#34;&gt;&lt;/i&gt;&#xA;    &lt;span class=&#34;sr-only&#34;&gt;Link to heading&lt;/span&gt;&#xA;  &lt;/a&gt;&#xA;&lt;/h2&gt;&#xA;&lt;figure class=&#34;estack&#34;&gt;&#xA;  &lt;style&gt;&#xA;    .estack { --es-hi: #2a78d6; --es-box: rgba(128,128,128,.10); --es-edge: rgba(128,128,128,.55);&#xA;              border: 1px solid rgba(128,128,128,.35); border-radius: 8px; padding: 14px; margin: 1.5em 0; }&#xA;    @media (prefers-color-scheme: dark) { body.colorscheme-auto .estack { --es-hi: #3987e5; } }&#xA;    body.colorscheme-dark .estack { --es-hi: #3987e5; }&#xA;    .estack svg { width: 100%; height: auto; display: block; }&#xA;    .estack svg text { fill: currentColor; font-size: 12.5px; }&#xA;    .estack svg .t { font-weight: 700; font-size: 13px; }&#xA;    .estack svg .s { font-size: 11px; opacity: .8; }&#xA;    .estack svg rect { fill: var(--es-box); stroke: var(--es-edge); }&#xA;    .estack svg rect.hi { stroke: var(--es-hi); stroke-width: 2; }&#xA;    .estack svg rect.ghost { fill: none; stroke-dasharray: 5 4; }&#xA;    .estack svg line, .estack svg path { stroke: var(--es-edge); stroke-width: 1.5; fill: none; }&#xA;    .estack figcaption { font-size: 12.5px; opacity: .8; margin-top: 8px; }&#xA;  &lt;/style&gt;&#xA;  &lt;svg viewBox=&#34;0 0 760 350&#34; role=&#34;img&#34; aria-label=&#34;Layers of an LLM serving stack&#34;&gt;&#xA;    &lt;defs&gt;&lt;marker id=&#34;es-arr&#34; viewBox=&#34;0 0 10 10&#34; refX=&#34;9&#34; refY=&#34;5&#34; markerWidth=&#34;7&#34; markerHeight=&#34;7&#34; orient=&#34;auto&#34;&gt;&lt;path d=&#34;M0,0 L10,5 L0,10 z&#34; style=&#34;fill:rgba(128,128,128,.7);stroke:none&#34;/&gt;&lt;/marker&gt;&lt;/defs&gt;&#xA;    &lt;rect x=&#34;270&#34; y=&#34;8&#34; width=&#34;480&#34; height=&#34;40&#34; rx=&#34;6&#34;/&gt;&#xA;    &lt;text class=&#34;t&#34; x=&#34;510&#34; y=&#34;33&#34; text-anchor=&#34;middle&#34;&gt;Your app: sends /v1/chat/completions&lt;/text&gt;&#xA;&#xA;    &lt;rect class=&#34;ghost&#34; x=&#34;10&#34; y=&#34;70&#34; width=&#34;240&#34; height=&#34;270&#34; rx=&#34;6&#34;/&gt;&#xA;    &lt;text class=&#34;t&#34; x=&#34;130&#34; y=&#34;98&#34; text-anchor=&#34;middle&#34;&gt;Hosted API&lt;/text&gt;&#xA;    &lt;text class=&#34;s&#34; x=&#34;130&#34; y=&#34;118&#34; text-anchor=&#34;middle&#34;&gt;Fireworks, Together, ...&lt;/text&gt;&#xA;    &lt;text class=&#34;s&#34; x=&#34;130&#34; y=&#34;190&#34; text-anchor=&#34;middle&#34;&gt;Runs every layer to the right&lt;/text&gt;&#xA;    &lt;text class=&#34;s&#34; x=&#34;130&#34; y=&#34;206&#34; text-anchor=&#34;middle&#34;&gt;for you. You see the API,&lt;/text&gt;&#xA;    &lt;text class=&#34;s&#34; x=&#34;130&#34; y=&#34;222&#34; text-anchor=&#34;middle&#34;&gt;not the engine, kernels or GPUs.&lt;/text&gt;&#xA;&#xA;    &lt;rect x=&#34;270&#34; y=&#34;70&#34; width=&#34;480&#34; height=&#34;56&#34; rx=&#34;6&#34;/&gt;&#xA;    &lt;text class=&#34;t&#34; x=&#34;285&#34; y=&#34;93&#34;&gt;Orchestration&lt;/text&gt;&#xA;    &lt;text class=&#34;s&#34; x=&#34;285&#34; y=&#34;112&#34;&gt;Dynamo, llm-d, Ray Serve: routing, autoscaling, multi-node&lt;/text&gt;&#xA;&#xA;    &lt;rect class=&#34;hi&#34; x=&#34;270&#34; y=&#34;138&#34; width=&#34;480&#34; height=&#34;56&#34; rx=&#34;6&#34;/&gt;&#xA;    &lt;text class=&#34;t&#34; x=&#34;285&#34; y=&#34;161&#34;&gt;Engine&lt;/text&gt;&#xA;    &lt;text class=&#34;s&#34; x=&#34;285&#34; y=&#34;180&#34;&gt;vLLM, SGLang, TensorRT-LLM: batching, KV cache, scheduling&lt;/text&gt;&#xA;&#xA;    &lt;rect x=&#34;270&#34; y=&#34;206&#34; width=&#34;480&#34; height=&#34;56&#34; rx=&#34;6&#34;/&gt;&#xA;    &lt;text class=&#34;t&#34; x=&#34;285&#34; y=&#34;229&#34;&gt;Kernels&lt;/text&gt;&#xA;    &lt;text class=&#34;s&#34; x=&#34;285&#34; y=&#34;248&#34;&gt;FlashAttention, FlashInfer, vendor kernels: GPU code&lt;/text&gt;&#xA;&#xA;    &lt;rect x=&#34;270&#34; y=&#34;274&#34; width=&#34;480&#34; height=&#34;56&#34; rx=&#34;6&#34;/&gt;&#xA;    &lt;text class=&#34;t&#34; x=&#34;285&#34; y=&#34;297&#34;&gt;Hardware&lt;/text&gt;&#xA;    &lt;text class=&#34;s&#34; x=&#34;285&#34; y=&#34;316&#34;&gt;NVIDIA, AMD, TPU, ... (or custom chips: Groq, Cerebras)&lt;/text&gt;&#xA;&#xA;    &lt;path d=&#34;M270,28 L130,28 L130,68&#34; marker-end=&#34;url(#es-arr)&#34;/&gt;&#xA;    &lt;line x1=&#34;510&#34; y1=&#34;48&#34; x2=&#34;510&#34; y2=&#34;68&#34; marker-end=&#34;url(#es-arr)&#34;/&gt;&#xA;  &lt;/svg&gt;&#xA;  &lt;figcaption&gt;The layers of a serving stack. Run it yourself and you pick each layer. Use a hosted API and the provider picks all of them.&lt;/figcaption&gt;&#xA;&lt;/figure&gt;&#xA;&#xA;&lt;p&gt;A replica is one running copy of a model, on as many GPUs as that copy needs, which mostly comes down to memory. The engine runs one replica: it takes requests, batches them, and runs them on the GPUs. A service runs several replicas behind a router when it needs more traffic than one copy can serve, or has to survive a failure. &lt;a href=&#34;https://hiren.me/posts/inference-engines-part-2/#how-many-replicas&#34; &gt;Part 2&lt;/a&gt; shows how to pick the number.&lt;/p&gt;</description>
    </item>
  </channel>
</rss>
