I work on LLM inference performance: GPU kernels, MoE tuning, and deterministic (batch-invariant) inference. I contribute to vLLM — mostly kernel-level performance work, measured on the hardware people actually deploy on.
Posts
-
Cutting the Determinism Tax: Making vLLM's Batch-Invariant Mode up to 1.4× Faster
How a two-GPU measurement project turned into my first merged vLLM PR (#53247) — and what it taught me about measurement discipline.
subscribe via RSS