I am now bridging the gap between LLM inference systems and their "Speed-of-Light", much as I once bridged the performance gap between CPU simulators and real silicon.
- I am a deep learning performance architect in NVIDIA, working on LLM inference system analysis, benchmarking, and modeling.
- I am systematically studying performance of emerging models and agentic serving system on cutting-edge hardwares.
- I participated in optimization and projection of performance on InferenceX V2 and AgentPerf for Blackwell.
- Before joining NVIDIA, I was a CPU performance architect in Beijing Institute of Open Source Chip (BOSC)
- I helped XiangShan CPU to achieve 15/GHz SPECint 2k6 score.
- I built a GEM5-based perf simulator for XiangShan CPU.





