Skip to content
View shinezyy's full-sized avatar

Block or report shinezyy

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
shinezyy/README.md

Hi there 👋

I am now bridging the gap between LLM inference systems and their "Speed-of-Light", much as I once bridged the performance gap between CPU simulators and real silicon.

  • I am a deep learning performance architect in NVIDIA, working on LLM inference system analysis, benchmarking, and modeling.
    • I am systematically studying performance of emerging models and agentic serving system on cutting-edge hardwares.
    • I participated in optimization and projection of performance on InferenceX V2 and AgentPerf for Blackwell.
  • Before joining NVIDIA, I was a CPU performance architect in Beijing Institute of Open Source Chip (BOSC)

Pinned Loading

  1. OpenXiangShan/GEM5 OpenXiangShan/GEM5 Public

    C++ 154 86

  2. micro-arch-training micro-arch-training Public

    How to make undergraduates or new graduates ready for advanced computer architecture research or modern CPU design

    665 48

  3. OpenXiangShan/NEMU OpenXiangShan/NEMU Public

    Super fast RISC-V ISA emulator for XiangShan processor

    C 341 142