我做 LLM 推理性能优化:GPU kernel、MoE 调优、确定性(batch-invariant)推理。我是 vLLM 贡献者——主要做 kernel 层的性能工作,在大家真正部署的硬件上做测量。
博客文章目前只有英文版。
文章
-
Cutting the Determinism Tax: Making vLLM's Batch-Invariant Mode up to 1.4× Faster
How a two-GPU measurement project turned into my first merged vLLM PR (#53247) — and what it taught me about measurement discipline.
subscribe via RSS