<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://lioeinaudi.github.io/feed.xml" rel="self" type="application/atom+xml" /><link href="https://lioeinaudi.github.io/" rel="alternate" type="text/html" /><updated>2026-08-31T03:16:58+00:00</updated><id>https://lioeinaudi.github.io/feed.xml</id><title type="html">Lio Einaudi</title><subtitle>Notes on LLM inference performance — kernels, determinism, and measurement discipline.</subtitle><entry><title type="html">Cutting the Determinism Tax: Making vLLM’s Batch-Invariant Mode up to 1.4× Faster</title><link href="https://lioeinaudi.github.io/2026/08/30/cutting-the-determinism-tax.html" rel="alternate" type="text/html" title="Cutting the Determinism Tax: Making vLLM’s Batch-Invariant Mode up to 1.4× Faster" /><published>2026-08-30T00:00:00+00:00</published><updated>2026-08-30T00:00:00+00:00</updated><id>https://lioeinaudi.github.io/2026/08/30/cutting-the-determinism-tax</id><content type="html" xml:base="https://lioeinaudi.github.io/2026/08/30/cutting-the-determinism-tax.html"><![CDATA[<p><em>How a two-GPU measurement project turned into my first merged vLLM PR (<a href="https://github.com/vllm-project/vllm/pull/53247">#53247</a>) — and what it taught me about measurement discipline.</em></p>

<h2 id="why-batch-invariance">Why batch invariance</h2>

<p>Ask an LLM served by any modern inference engine the same question twice — greedy sampling, temperature 0 — and you can get two different answers. Not because of sampling: because your request was batched with different neighbors each time. Reduction order inside the kernels changes with batch composition, floating-point addition is not associative, and logprobs drift by a few ULPs. Usually nobody cares. But if you are doing RL rollouts and want bit-exact replay, or debugging a training-inference mismatch, this non-determinism is poison.</p>

<p>vLLM ships a batch-invariant mode (<code class="language-plaintext highlighter-rouge">VLLM_BATCH_INVARIANT=1</code>) that swaps the offending kernels for invariant ones. The catch: those kernels are slow. The community calls the slowdown the <strong>determinism tax</strong>, and when I started measuring it, the tax on my hardware was 30–35% of throughput.</p>

<p>This post is about finding out exactly where that tax is paid and tuning most of the biggest line-item away. The work landed upstream as <a href="https://github.com/vllm-project/vllm/pull/53247">PR #53247</a>; the headline numbers on H20 are <strong>−29% end-to-end latency (1.42×) at batch 8</strong> and <strong>+3–12% throughput</strong>, with the invariance guarantee bit-for-bit intact.</p>

<h2 id="setup-two-gpus-the-big-labs-dont-tune-for">Setup: two GPUs the big labs don’t tune for</h2>

<p>I rented two kinds of cloud instances deliberately outside the H100/B200 mainstream:</p>

<ul>
  <li><strong>NVIDIA H20</strong> — 96 GB, Hopper (sm90), huge bandwidth but ~1/6 of H100’s compute. The dominant inference card in Chinese production fleets.</li>
  <li><strong>RTX 4090D</strong> — 24 GB, Ada (sm89). Consumer silicon, no TMA, no WGMMA.</li>
</ul>

<p>Model: Qwen3-1.7B. Two harnesses:</p>

<ol>
  <li><strong>A tax benchmark</strong>: serving throughput and latency with <code class="language-plaintext highlighter-rouge">VLLM_BATCH_INVARIANT</code> on/off, compiled and eager.</li>
  <li><strong>An invariance probe</strong>: run a “needle” prompt alone, then mixed into a batch of 32 (greedy, prefix cache off), and compare logprobs <strong>bit-for-bit in fp32</strong>. A tuned kernel only counts if the probe stays green.</li>
</ol>

<p>The probe matters as much as the benchmark. Every performance claim below passed it.</p>

<h2 id="first-lesson-before-any-tuning-check-your-version">First lesson before any tuning: check your version</h2>

<p>My first run showed something alarming: batch-invariant mode <em>failing its own guarantee</em> in eager mode on both GPUs — prefill logprobs off by ~0.02. A real upstream bug? Almost: it was a real bug that had <strong>already been fixed</strong>. The root cause (a C++ RMSNorm kernel whose reduction block size changes with token count, which batch-invariant mode couldn’t intercept) was fixed upstream weeks earlier — but I had benchmarked a stale release that predated the fix. On the current release, eager went green.</p>

<p>Cheap lesson, permanently installed: <strong>benchmarks run on the latest release or they don’t count.</strong> Version capture went into the checklist next to GPU model and driver.</p>

<p>Re-measured on the current release, the true tax (compiled, throughput): <strong>29.9% on 4090D, 35.0% on H20</strong>. Latency tax was worse: +73% and +110%.</p>

<h2 id="finding-where-the-tax-is-paid">Finding where the tax is paid</h2>

<p>The batch-invariant module has per-architecture code paths that <em>look</em> load-bearing — a cuBLAS workspace configuration for Hopper, aten overrides for sm8x. I profiled instead of trusting the code.</p>

<p>The profile said something simpler: with invariance on, every cuBLASLt GEMM disappears, replaced by a single Triton kernel — <code class="language-plaintext highlighter-rouge">matmul_kernel_persistent</code> — because vLLM’s linear layer routes <strong>all architectures</strong> straight into it. In one H20 profile it was called 14,577 times, accounted for <strong>79.4% of batch-invariant kernel time</strong>, and explained ~97% of the total kernel-time increase. The per-arch branches I’d been reading were dead code for this path. (A workspace-size ablation confirmed it: varying it changed end-to-end time by ≤0.23%.)</p>

<p>One kernel, hardcoded launch configs, all architectures, all shapes. That’s not a tax, that’s a tuning opportunity.</p>

<h2 id="the-constraint-that-makes-tuning-legal">The constraint that makes tuning legal</h2>

<p>You can’t just autotune an invariant kernel — tuning is exactly how you break invariance. The way out is to look at <em>why</em> the kernel is invariant: each output row’s value depends only on the <strong>K-loop reduction order</strong>. So:</p>

<blockquote>
  <p><strong><code class="language-plaintext highlighter-rouge">BLOCK_K</code> must stay constant across all M for a fixed weight shape (N, K).</strong>
<code class="language-plaintext highlighter-rouge">BLOCK_M</code>, <code class="language-plaintext highlighter-rouge">BLOCK_N</code>, <code class="language-plaintext highlighter-rouge">num_warps</code>, <code class="language-plaintext highlighter-rouge">num_stages</code> are free to vary per M-bucket — they change scheduling, not reduction order.</p>
</blockquote>

<p>That single invariant separates the tunable parameters from the untouchable one, and it’s the technical core of the PR. Every candidate config still had to pass the bitwise probe — trust the argument, verify the bits.</p>

<p>A pleasant side-find from the correctness harness: at K=2048 on the 4090D, the batch-invariant Triton kernel <em>passed</em> an fp32 reference check that <code class="language-plaintext highlighter-rouge">torch.mm</code>’s default path failed — PyTorch’s bf16 reduced-precision reduction accumulates enough error to lose to the “slow deterministic” kernel on accuracy. Deterministic and <em>more</em> accurate.</p>

<h2 id="the-sweep">The sweep</h2>

<p>A 19,200-point sweep across both GPUs — M-buckets × <code class="language-plaintext highlighter-rouge">BLOCK_M</code>/<code class="language-plaintext highlighter-rouge">BLOCK_N</code>/<code class="language-plaintext highlighter-rouge">num_warps</code>/<code class="language-plaintext highlighter-rouge">num_stages</code>, <code class="language-plaintext highlighter-rouge">BLOCK_K</code> pinned per shape, two correctness gates per config. Results:</p>

<table>
  <thead>
    <tr>
      <th> </th>
      <th>decode GEMMs (vs torch)</th>
      <th>prefill</th>
      <th>E2E throughput tax</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>RTX 4090D</td>
      <td>0.25× → <strong>0.77×</strong> (3.0×)</td>
      <td>1.16×</td>
      <td>29.9% → <strong>21.9%</strong></td>
    </tr>
    <tr>
      <td>H20</td>
      <td>0.21× → <strong>0.59×</strong> (2.8×)</td>
      <td>1.11×</td>
      <td>35.0% → <strong>31.0%</strong></td>
    </tr>
  </tbody>
</table>

<p>Latency tax dropped harder: 73→36% (4090D), 110→59% (H20).</p>

<p>Two measurement stories from this phase that I’d repeat on any project:</p>

<ul>
  <li><strong>The eager “regression” that wasn’t.</strong> Tuned eager throughput initially looked 44% <em>worse</em> on 4090D. Cause: M-bucketing multiplies Triton JIT compilations (3 → 43 cache entries), and a short benchmark eats the cold-start. Steady-state after warmup: +1.2%. Compiled mode prepays this in graph capture, which is why it never showed there.</li>
  <li><strong>The phantom H20 number.</strong> One H20 eager baseline was ~40% slower than every other run. Before explaining it, I checked node identity — GPU UUID and serial numbers showed the before/after had landed on <em>different physical machines</em>. Re-ran both sides on one node, with the serials logged: the anomaly vanished, and one number I had already posted publicly needed a correction. <strong>Before/after on provably the same node, or it didn’t happen.</strong></li>
</ul>

<h2 id="two-ambushes-from-torchcompile">Two ambushes from torch.compile</h2>

<p>The config-lookup code sits on a hot path that <code class="language-plaintext highlighter-rouge">torch.compile</code> traces. Two “clean” implementations died there:</p>

<ol>
  <li><code class="language-plaintext highlighter-rouge">functools.lru_cache</code> around device-name lookup → Dynamo traces <em>through</em> the cache into pynvml’s ctypes calls → <code class="language-plaintext highlighter-rouge">Unsupported</code> → <strong>the compiled engine doesn’t start.</strong> Kernel tests and eager pytest were all green; only a real compiled end-to-end run caught it. Shipping it would have been a P0.</li>
  <li><code class="language-plaintext highlighter-rouge">bisect_left(..., key=...)</code> → also <code class="language-plaintext highlighter-rouge">Unsupported</code> in the current torch.</li>
</ol>

<p>Final shape: resolve the device to a config table <strong>once at init</strong>, store it in a module-level variable, and look up buckets with a linear scan over a sorted tuple of ≤10 entries. Boring, trace-proof, and just as fast.</p>

<p>New rule in my acceptance checklist: <strong>if code can be traced by Dynamo, the test plan includes a real compiled end-to-end run.</strong> No proxy suffices.</p>

<h2 id="the-review-arc">The review arc</h2>

<p>I opened the PR keyed to exact device names, with every one of the 100 config entries mechanically cross-checked against the sweep JSON. First maintainer response, same day:</p>

<blockquote>
  <p>“don’t want to tune triton for older architecture, not used frequently in production”</p>
</blockquote>

<p>Fair concern, wrong premise — the H20 is not an older architecture; it’s current-generation Hopper, and it <em>is</em> production hardware for a large fraction of real deployments. I replied with that one factual clarification (no argument beyond it) plus an offer to narrow scope. The maintainer’s answer: “Ohh make sense for hopper — could you test e2e using vllm bench? Let’s see if it’s worth the complexity.”</p>

<p>So the decision criterion became end-to-end numbers in the exact format of the maintainer’s own prior perf PR — command lines and raw before/after output blocks, same-node discipline, upstream’s full determinism test suite passing:</p>

<ul>
  <li><strong>H20</strong>: latency −29.4% at batch 8 (1.42×), −19.0% at batch 32; throughput +3.4%</li>
  <li><strong>RTX 4090D</strong>: latency −23.0% at batch 8, −16.5% at batch 32; throughput +12.0%</li>
</ul>

<p>The review flipped:</p>

<blockquote>
  <p>“Wow, very good perf improvement I didn’t think about before, nice work”</p>
</blockquote>

<p>Three review rounds later (configs moved to their own file, keys widened from exact device names to architecture families — <code class="language-plaintext highlighter-rouge">hopper</code> backed by H20 measurements, <code class="language-plaintext highlighter-rouge">ada</code> by 4090D — as the maintainer requested), the PR merged. <strong>Seven days from first measurement to merge.</strong> The maintainer plans follow-up tuning and refactoring on top, and the architecture-family keys leave labeled slots for Blackwell when someone with the hardware shows up.</p>

<p>One more moment worth recording: mid-review, another contributor offered to add H100 configs to the PR — derived numbers, not measured ones. I asked them to hold for a follow-up PR after real measurements instead. The entire credibility of a tuning PR is that <em>every number has a provenance</em>; one estimated row poisons the table.</p>

<h2 id="what-generalized">What generalized</h2>

<p>The kernel work is the resume line, but the transferable part is the discipline:</p>

<ol>
  <li><strong>Latest release or it doesn’t count.</strong> My scariest “bug” was a stale version.</li>
  <li><strong>Profile before believing code structure.</strong> The per-arch branches were dead; one kernel was 79% of the story.</li>
  <li><strong>Find the invariant that makes optimization legal.</strong> <code class="language-plaintext highlighter-rouge">BLOCK_K</code> constancy turned “you can’t tune a deterministic kernel” into a 19,200-point sweep.</li>
  <li><strong>Same node, logged serials, or re-run.</strong> One cross-node number cost me a public correction.</li>
  <li><strong>Compiled E2E is a non-negotiable gate</strong> for anything Dynamo might trace.</li>
  <li><strong>In review, one factual clarification beats a debate</strong> — and the reviewer’s own preferred evidence format is the fastest path to yes.</li>
  <li><strong>Every number needs provenance</strong> — in your tables and everyone else’s.</li>
</ol>

<hr />

<p><em>The tax isn’t zero yet: decode GEMMs still trail cuBLAS, and the remaining gap lives in shapes and architectures nobody has swept. If you have the hardware, the config table has labeled empty slots.</em></p>]]></content><author><name></name></author><summary type="html"><![CDATA[How a two-GPU measurement project turned into my first merged vLLM PR (#53247) — and what it taught me about measurement discipline.]]></summary></entry></feed>