Skip to content
View andyxning's full-sized avatar
🎯
Focusing
🎯
Focusing
  • Beijing, China

Organizations

@nsqio @kubernetes

Block or report andyxning

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Starred repositories

Showing results

Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing,…

Rust 455 134 Updated Aug 10, 2026

Recipes and resources for training, building generative AI with Fireworks

Jupyter Notebook 189 53 Updated Aug 10, 2026

Model compression toolkit engineered for enhanced usability, comprehensiveness, and efficiency.

Python 1,516 173 Updated Aug 7, 2026

DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms

Python 6,915 646 Updated Jul 9, 2026

DeepEP: an efficient expert-parallel communication library

Cuda 9,971 1,370 Updated Aug 5, 2026

mKernel: fast multi-node, multi-GPU fused kernels

Cuda 264 24 Updated Aug 10, 2026

Tile primitives for speedy kernels

Cuda 3,621 316 Updated Jul 13, 2026

A curated list of best cuda programming books

947 31 Updated May 19, 2026

Machine Learning Engineering Open Book

Python 18,576 1,197 Updated Aug 8, 2026

Module, Model, and Tensor Serialization/Deserialization

Python 320 54 Updated Jul 7, 2026

Benchmark suite for LLMs from Fireworks.ai

Python 110 39 Updated Aug 6, 2026

High Performance LLM Inference Operator Library

C++ 1,100 128 Updated Aug 6, 2026

📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉

Python 5,450 431 Updated Jul 26, 2026

SkyRL: A Modular Full-stack RL Library for LLMs

Python 2,140 400 Updated Aug 6, 2026

LLMPerf is a library for validating and benchmarking LLMs

Python 1,129 203 Updated Dec 9, 2024

Manages Unified Access to Generative AI Services built on Envoy Gateway

Go 1,914 326 Updated Aug 10, 2026
C++ 548 46 Updated Jul 14, 2026

Inference server benchmarking tool

Rust 168 32 Updated Jun 9, 2026

Fluid, elastic data abstraction and acceleration for BigData/AI applications in cloud. (Project under CNCF)

Go 1,961 1,274 Updated Aug 7, 2026

Using CRDs to manage GPU resources in Kubernetes.

Go 214 30 Updated Nov 21, 2022

Heterogeneous GPU Sharing on Kubernetes

Go 4,286 749 Updated Aug 10, 2026

AI on GKE is a collection of examples, best-practices, and prebuilt solutions to help build, deploy, and scale AI Platforms on Google Kubernetes Engine

Jupyter Notebook 330 242 Updated Jun 23, 2025

Cloud Native Benchmarking of Foundation Models

Python 46 20 Updated Jul 31, 2025

A fast GPU memory copy library based on NVIDIA GPUDirect RDMA technology

C 1,404 193 Updated Jul 14, 2026

LLM KV cache compression made easy

Python 1,164 168 Updated Aug 10, 2026

Open source AI coding agent. Designed for large projects and real world tasks.

Go 15,581 1,170 Updated Oct 3, 2025

A CLI inspector for the Model Context Protocol

JavaScript 442 41 Updated Jun 8, 2026

Serving multiple LoRA finetuned LLM as one

Python 1,171 65 Updated May 8, 2024
Next