<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Rollout on </title>
    <link>https://hiren.me/tags/rollout/</link>
    <description>Recent content in Rollout on </description>
    <generator>Hugo</generator>
    <language>en</language>
    <lastBuildDate>Wed, 30 Sep 2026 12:30:00 -0700</lastBuildDate>
    <atom:link href="https://hiren.me/tags/rollout/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Inference in production (part 4): rolling out a new model</title>
      <link>https://hiren.me/posts/inference-in-production-part-4/</link>
      <pubDate>Wed, 30 Sep 2026 12:30:00 -0700</pubDate>
      <guid>https://hiren.me/posts/inference-in-production-part-4/</guid>
      <description>&lt;p&gt;The first three parts followed a request through &lt;a href=&#34;https://hiren.me/posts/inference-in-production-part-1/&#34; &gt;prefill and decode pools&lt;/a&gt;, &lt;a href=&#34;https://hiren.me/posts/inference-in-production-part-2/&#34; &gt;the router and autoscaler in front of them&lt;/a&gt;, and &lt;a href=&#34;https://hiren.me/posts/inference-in-production-part-3/&#34; &gt;what happens when the hardware under them fails&lt;/a&gt;. This part is about changing what runs on all of it: a new model, a new checkpoint of the same model, or a new engine version.&lt;/p&gt;&#xA;&lt;p&gt;A rollout does to every replica, on purpose, what the last two posts measured happening by accident. Each old replica is shut down, like part 3&amp;rsquo;s failures, and each new one pays &lt;a href=&#34;https://hiren.me/posts/inference-in-production-part-2/#why-a-new-replica-takes-minutes&#34; &gt;part 2&amp;rsquo;s cold start&lt;/a&gt;. A rollout goes well when users don&amp;rsquo;t notice either. The sources are project docs and operators&amp;rsquo; own write-ups; the part about draining I checked on Modal.&lt;/p&gt;</description>
    </item>
  </channel>
</rss>
