- Lead the Machine Learning Infrastructure team that powers Xiaohongshu (小红书) and RedNote's core Search, Ads, and Recommendation systems, serving as both technical lead and hands-on core contributor.
- Founded and currently lead the AI Infrastructure team (AIGC / LLM / VLM); architected and delivered a production-grade unified framework spanning pretraining, post-training, and inference.
- Built a highly performance-optimized LLM serving framework — deployed in production Agent scenarios such as Xiaohongshu AI Search (AI搜索) and Diandian (点点), serving 300 million MAU.
- Lead the team driving AIGC inference optimization, pushing toward speed-of-light performance across GPU, NPU, PPU, and XPU platforms.
Academic Publications 10 papers · Google Scholar ↗
Awards 1
- 🏆 2025 Xiaohongshu Impact Challenge — Business Breakthrough Award, Annual Champion
Press & Media Coverage 9 articles
Open Source & Community 4
LLM ServingAIGCVLM
Unified Training FrameworkGPU · NPU · XPU
Tencent · 腾讯 — Deep Learning Framework Senior Staff
May 2020 — Aug 2025 · 5 yrs 4 mos · Shenzhen
- Founded the open-source LLM inference engine KsanaLLM (一念) (github.com/Tencent/KsanaLLM) as main contributor — throughput 45% higher than SOTA frameworks SGLang & vLLM on DeepSeek V3/R1 (benchmark 2025-06-13), deployed for ima.copilot, QQ AI Agents, etc.
- Developed AI infrastructure for Tencent's Hunyuan (混元) LLM across NVIDIA GPU, Huawei Ascend (昇腾) NPU, and Enflame (燧原) / ZiXiao (紫霄) GCU.
- Co-founded the Venus AI Draw Engine — VLM & diffusion (video/image generation) inference acceleration for Tencent Zenvideo (智影) and game art assets (Honor of Kings 王者荣耀, PUBG Mobile 和平精英).
- Co-founded Numerous (无量), PCG's heterogeneous ML platform for large-scale recommender training & inference across QQ, Tencent Video & QQ Music — 10B+ queries/day; main architect of GameLoop's (gameloop.com) international federated learning framework.
- Contributed to NVIDIA HugeCTR — co-developed with the NVIDIA team, broke MLPerf training world records in 2020 & 2021 (NVIDIA Developer Blog); also contributed to S-LoRA (PR #13).
- Served as PCG C++ Committee Member (Nov 2021 – Aug 2025): Tencent engineer promotion assessor; participated in planning Tencent's technical blueprint in AI/ML; code contributor to company-level infrastructure.
- Granted patent CN118446316A.
Academic Publications 1
- Shen, W., Liu, Z., Tan, Y., Luo, Z. & Lei, Z. (2023). KubeGPU: efficient sharing and isolation mechanisms for GPU resource management in container cloud. The Journal of Supercomputing, 79(1), pp. 591–625.
Awards 6
- 🥇 2024 Tencent PCG H1 CVP Technology Innovation Award
- 🥉 2024 Tencent Technology Breakthrough Award — Bronze Prize
- ⭐ Tencent Excellent Employee — 2020 / 2021 / 2022 / 2023
- ⭐ Tencent Outstanding Contributor — 2020 / 2022 / 2023
- ⭐ Excellent Contributor, Tencent Open Source Collaboration Project — 2021 / 2022 / 2023
- ⭐ 2020 Tencent Excellence in R&D Award · 2020 Tencent Open Source Collaboration Award
KsanaLLMHunyuanHugeCTR
DiffusionRecSysFederated Learning
Alibaba Group · 阿里巴巴 — Deep Learning Platform Engineer
Oct 2018 — May 2020 · 1 yr 8 mos · Hangzhou
DAMO Academy (达摩院) · Autonomous Driving Lab (自动驾驶实验室) · Automated Logistics Transportation Project (自动驾驶物流车项目 · 小蛮驴)
- Optimized the model-training pipeline for autonomous driving — graph optimization, neural-network pruning, deep compression & quantization, knowledge distillation, and hardware-aware Neural Architecture Search (NAS) with AutoML-Keras.
- Accelerated inference via TVM operator tuning and custom operator development for NVIDIA GPU & Huawei NPU; built the edge inference engine on Autopilot for the Xiaomanlv (小蛮驴) autonomous logistics vehicle.
TVMModel CompressionNASEdge Inference
NVIDIA — Deep Learning Software Engineer
May 2016 — Oct 2018 · 2 yrs 6 mos · Shanghai
- Developed part of TensorRT features (release 4.0+, RC 5.0+) and tuned its inference performance across different models on DrivePX2 and Jetson TX2/TX1 platforms (Python/C++), including weight quantization & calibration (INT8, INT4).
- Tuned cuDNN performance on DrivePX2 / Jetson TX2/TX1 platforms with Python/C++.
- Built automation projects GFE, GFN and COSMOS.
TensorRTcuDNNCUDAINT8 / INT4
Ele.me · 饿了么 — Engineer Intern
Sep 2015 — Mar 2016 · Shanghai
- Developed various Talaris system APIs on the Vespone combining framework: Flask + Vespone + Thrift + Redis + MySQL.
Intel — Engineer Intern
May 2014 — Aug 2015 · 1 yr 4 mos
- Automated the profiling and performance tuning of the Intel Xeon Phi many-core processor (PVL team).
- Built an exclusive distributed automated testing system based on MCG's auto-testing framework — a 3-level architecture: Django (deployed on UWSGI + Nginx for user interaction) + Celery (scheduling tasks across hosts) + testing framework, with Redis cache and MySQL persistence.