<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>InfiniBand on </title>
    <link>https://hiren.me/tags/infiniband/</link>
    <description>Recent content in InfiniBand on </description>
    <generator>Hugo</generator>
    <language>en</language>
    <lastBuildDate>Tue, 29 Sep 2026 11:30:00 -0700</lastBuildDate>
    <atom:link href="https://hiren.me/tags/infiniband/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Inference in production (part 1): disaggregated inference on B200 and GB200 NVL72</title>
      <link>https://hiren.me/posts/inference-in-production-part-1/</link>
      <pubDate>Tue, 29 Sep 2026 11:30:00 -0700</pubDate>
      <guid>https://hiren.me/posts/inference-in-production-part-1/</guid>
      <description>&lt;p&gt;This series is about how LLM inference is deployed in production: what the hardware looks like, where each piece of work runs, and how one request moves through it. The numbers in this post come from NVIDIA&amp;rsquo;s documentation and from what DeepSeek and Moonshot have published about their own systems. Where a piece overlaps with something I measured earlier, I link to it.&lt;/p&gt;&#xA;&lt;p&gt;The short version: production systems split each request&amp;rsquo;s two phases across separate pools of GPUs, and the design question that changes the picture most is where the NVLink boundary sits. On an HGX B200 system it&amp;rsquo;s the node, so the KV cache crosses the network between the pools. On a GB200 NVL72 it&amp;rsquo;s the whole rack, so the cache can stay on NVLink, as long as the deployment fits inside that rack.&lt;/p&gt;</description>
    </item>
  </channel>
</rss>
