Zhaode Wang
AI Inference Engine Expert · On-Device LLM · High-Performance Computing
Education
Institute of Computing Technology, Chinese Academy of Sciences State Key Laboratory of Computer Architecture
M.S. in Computer Architecture
Shandong University
B.S. in Computer Science and Technology
Experience
Alibaba · Taotian Group MNN Tech Lead
- MNN core architecture: Lead MNN architecture and heterogeneous performance optimization across ARM CPUs, Metal/OpenCL/Vulkan GPUs, and Hexagon/CoreML/NNAPI accelerators, covering operator kernels, memory planning, graph optimization and fusion, and model conversion.
- On-device AI platform: Built an embedded Python runtime and end-to-end deployment stack with CV/audio modules, Python APIs, and CI/CD, enabling dozens of algorithms for Pailitao visual search, Taobao Live, security, and feed.
- MNN-LLM: Initiated and lead MNN-LLM for 100+ language and multimodal models, with 2/3/4/8-bit and mixed-precision quantization, KV-cache and attention optimization, and speculative decoding.
- Production & open source: Drove MNN adoption across Taobao, Xianyu, Quark, and Qwen; deployed on-device LLMs in Taobao and Xianyu search and recommendation scenarios at tens-of-millions-user scale; supported DingTalk, Youku, Taobao Instant Commerce, and Amap; grew MNN to 15.8K+ GitHub stars.
Cambricon Technologies Compiler Engineer Intern
- Developed compiler and linker components for Cambricon's BANG toolchain, focusing on heterogeneous linking; subsequently optimized Caffe operators for MLU accelerators.
Megvii Algorithm Engineer Intern
- Optimized traffic video understanding algorithms and productionized them in C++.
Microsoft Asia Engineering Academy Software Engineer Intern
- Evaluated LevelDB, Presto, and Hadoop for advertising data storage and analytics workloads.
Selected Publications
Core Contributions
MM Asia
Collaborative Research
ACM MM
Patents & Applications
BANG-Linker: Heterogeneous Linking Method, Apparatus, and Related Products BANG Simulator: Programming Debugging Method, Apparatus, and Related Products
Cambricon · Filed Solution for Deploying Large Language Models on Mobile Devices
Alibaba · Filed Saliency-Aware Blockwise Quantization via Floating-Point Zero-Point Degrees of Freedom Mixed-Precision Allocation for Low-Bit LLMs Based on Action-Conditioned Operator Error
Alibaba · Pending Honors & Competitions
IEEE AICAS Grand Challenge: LLM Software and Hardware System Co-optimization
First Prize NPU Competitions: Tecorigin Operator Development & SOPHGO TPU Programming
First Prize Skills
Languages C, C++, Python, Assembly, CUDA, OpenCL, Vulkan, Metal, BANG, Java, Objective-C
Expertise Compilers, AI inference engines, CPU/GPU/NPU operator optimization, model conversion, quantization, large language models