I work on LLM inference performance: GPU kernels, MoE tuning, and deterministic (batch-invariant) inference. I contribute to vLLM — mostly kernel-level performance work, measured on the hardware people actually deploy on.

Posts

subscribe via RSS