- Fairfax, VA
- wangaoone.github.io
Stars
verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
Agent Reinforcement Trainer: train multi-step agents for real-world tasks using GRPO. Give your agents on-the-job training. Reinforcement learning for Qwen3.6, GPT-OSS, Llama, and more!
Post-training with Tinker
A next.js web application that integrates AI capabilities with draw.io diagrams. This app allows you to create, modify, and enhance diagrams through natural language commands and AI-assisted visual…
slime is an LLM post-training framework for RL Scaling.
OpenTinker is an RL-as-a-Service infrastructure for foundation models
Checkpoint-engine is a simple middleware to update model weights in LLM inference engines
TurboDiffusion: 100–200× Acceleration for Video Diffusion Models
A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.
EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework based on veRL
An Efficient and User-Friendly Scaling Library for Reinforcement Learning with Large Language Models
FaaSNet: Scalable and Fast Provisioning of Custom Serverless Container Runtimes at Alibaba Cloud Function Compute (USENIX ATC'21)
High-quality and streaming Speech-to-Speech interactive agent in a single file. 只用一个文件实现的流式全双工语音交互原型智能体!
A high-performance inference system for large language models, designed for production environments.
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
KV cache store for distributed LLM inference
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
A high-performance distributed file system designed to address the challenges of AI training and inference workloads.
Production-tested AI infrastructure tools for efficient AGI development and community-driven innovation
High-performance Python librarys for connecting AI/ML frameworks with OSS storage.
主要记录大语言大模型(LLMs) 算法(应用)工程师相关的知识及面试题
A self-learning tutorail for CUDA High Performance Programing.
Disaggregated serving system for Large Language Models (LLMs).
A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations
SGLang is a high-performance serving framework for large language models and multimodal models.
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. Tensor…