<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Kubernetes on </title>
    <link>https://hiren.me/tags/kubernetes/</link>
    <description>Recent content in Kubernetes on </description>
    <generator>Hugo</generator>
    <language>en</language>
    <lastBuildDate>Wed, 30 Sep 2026 12:30:00 -0700</lastBuildDate>
    <atom:link href="https://hiren.me/tags/kubernetes/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Inference in production (part 4): rolling out a new model</title>
      <link>https://hiren.me/posts/inference-in-production-part-4/</link>
      <pubDate>Wed, 30 Sep 2026 12:30:00 -0700</pubDate>
      <guid>https://hiren.me/posts/inference-in-production-part-4/</guid>
      <description>&lt;p&gt;The first three parts followed a request through &lt;a href=&#34;https://hiren.me/posts/inference-in-production-part-1/&#34; &gt;prefill and decode pools&lt;/a&gt;, &lt;a href=&#34;https://hiren.me/posts/inference-in-production-part-2/&#34; &gt;the router and autoscaler in front of them&lt;/a&gt;, and &lt;a href=&#34;https://hiren.me/posts/inference-in-production-part-3/&#34; &gt;what happens when the hardware under them fails&lt;/a&gt;. This part is about changing what runs on all of it: a new model, a new checkpoint of the same model, or a new engine version.&lt;/p&gt;&#xA;&lt;p&gt;A rollout does to every replica, on purpose, what the last two posts measured happening by accident. Each old replica is shut down, like &lt;a href=&#34;https://hiren.me/posts/inference-in-production-part-3/#what-happens-to-the-requests-in-flight&#34; &gt;part 3&amp;rsquo;s failures&lt;/a&gt;, and each new one pays &lt;a href=&#34;https://hiren.me/posts/inference-in-production-part-2/#why-a-new-replica-takes-minutes&#34; &gt;part 2&amp;rsquo;s cold start&lt;/a&gt;. A rollout goes well when users don&amp;rsquo;t notice either. The sources are project docs and operators&amp;rsquo; own write-ups; the part about draining I checked on Modal.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Inference in production (part 3): when hardware goes bad</title>
      <link>https://hiren.me/posts/inference-in-production-part-3/</link>
      <pubDate>Wed, 30 Sep 2026 11:00:00 -0700</pubDate>
      <guid>https://hiren.me/posts/inference-in-production-part-3/</guid>
      <description>&lt;p&gt;&lt;a href=&#34;https://hiren.me/posts/inference-in-production-part-1/&#34; &gt;Part 1&lt;/a&gt; followed a request through prefill and decode pools, and &lt;a href=&#34;https://hiren.me/posts/inference-in-production-part-2/&#34; &gt;part 2&lt;/a&gt; covered the gateway, router and autoscaler in front of them. This part is about what happens when a GPU, a node or an NVL72 tray fails under them: how often that happens, how it shows up, how much of the deployment it takes out, and what becomes of the requests that were running on it. The failure data comes from published papers and NVIDIA&amp;rsquo;s docs; the part about requests in flight I measured by killing vLLM replicas on Modal. The script is in the &lt;a href=&#34;https://github.com/hirenp/kv-cache-lab&#34;  class=&#34;external-link&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;lab repo&lt;/a&gt;.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Inference in production (part 2): the gateway, the router and the autoscaler</title>
      <link>https://hiren.me/posts/inference-in-production-part-2/</link>
      <pubDate>Wed, 30 Sep 2026 00:40:00 -0700</pubDate>
      <guid>https://hiren.me/posts/inference-in-production-part-2/</guid>
      <description>&lt;p&gt;&lt;a href=&#34;https://hiren.me/posts/inference-in-production-part-1/&#34; &gt;Part 1&lt;/a&gt; followed one request through prefill and decode pools on B200 and GB200 NVL72. This part covers what sits in front of those pools: the gateway that accepts the request, the router that picks the GPUs for it, and the autoscaler that decides how many GPUs there are. The last section measures why a new replica takes minutes to become useful, with vLLM cold starts I timed on H100s. The rest comes from project docs and published benchmarks, linked where they&amp;rsquo;re used.&lt;/p&gt;</description>
    </item>
  </channel>
</rss>
