zhaode bits & days

Zhaode Wang

AI Inference Engine Expert · On-Device LLM · High-Performance Computing

Education

Institute of Computing Technology, Chinese Academy of Sciences State Key Laboratory of Computer Architecture
M.S. in Computer Architecture
Shandong University
B.S. in Computer Science and Technology

Experience

Alibaba · Taotian Group MNN Tech Lead
  • MNN core architecture: Lead MNN architecture and heterogeneous performance optimization across ARM CPUs, Metal/OpenCL/Vulkan GPUs, and Hexagon/CoreML/NNAPI accelerators, covering operator kernels, memory planning, graph optimization and fusion, and model conversion.
  • On-device AI platform: Built an embedded Python runtime and end-to-end deployment stack with CV/audio modules, Python APIs, and CI/CD, enabling dozens of algorithms for Pailitao visual search, Taobao Live, security, and feed.
  • MNN-LLM: Initiated and lead MNN-LLM for 100+ language and multimodal models, with 2/3/4/8-bit and mixed-precision quantization, KV-cache and attention optimization, and speculative decoding.
  • Production & open source: Drove MNN adoption across Taobao, Xianyu, Quark, and Qwen; deployed on-device LLMs in Taobao and Xianyu search and recommendation scenarios at tens-of-millions-user scale; supported DingTalk, Youku, Taobao Instant Commerce, and Amap; grew MNN to 15.8K+ GitHub stars.
Cambricon Technologies Compiler Engineer Intern
  • Developed compiler and linker components for Cambricon's BANG toolchain, focusing on heterogeneous linking; subsequently optimized Caffe operators for MLU accelerators.
Megvii Algorithm Engineer Intern
  • Optimized traffic video understanding algorithms and productionized them in C++.
Microsoft Asia Engineering Academy Software Engineer Intern
  • Evaluated LevelDB, Presto, and Hadoop for advertising data storage and analytics workloads.

Selected Publications

Patents & Applications

BANG-Linker: Heterogeneous Linking Method, Apparatus, and Related Products BANG Simulator: Programming Debugging Method, Apparatus, and Related Products
Cambricon · Filed
Solution for Deploying Large Language Models on Mobile Devices
Alibaba · Filed
Saliency-Aware Blockwise Quantization via Floating-Point Zero-Point Degrees of Freedom Mixed-Precision Allocation for Low-Bit LLMs Based on Action-Conditioned Operator Error
Alibaba · Pending

Honors & Competitions

IEEE AICAS Grand Challenge: LLM Software and Hardware System Co-optimization
First Prize
NPU Competitions: Tecorigin Operator Development & SOPHGO TPU Programming
First Prize

Skills

Languages C, C++, Python, Assembly, CUDA, OpenCL, Vulkan, Metal, BANG, Java, Objective-C
Expertise Compilers, AI inference engines, CPU/GPU/NPU operator optimization, model conversion, quantization, large language models