Skip to content
View andyxning's full-sized avatar
🎯
Focusing
🎯
Focusing
  • Beijing, China

Organizations

@nsqio @kubernetes

Block or report andyxning

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Starred repositories

Showing results

Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing,…

Rust 413 124 Updated Jul 24, 2026

Recipes and resources for building, deploying, and fine-tuning generative AI with Fireworks.

Jupyter Notebook 187 52 Updated Jul 25, 2026

Model compression toolkit engineered for enhanced usability, comprehensiveness, and efficiency.

Python 1,483 168 Updated Jul 23, 2026

DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms

Python 6,765 628 Updated Jul 9, 2026

DeepEP: an efficient expert-parallel communication library

Cuda 9,885 1,337 Updated Jul 14, 2026

mKernel: fast multi-node, multi-GPU fused kernels

Cuda 255 24 Updated Jun 21, 2026

Tile primitives for speedy kernels

Cuda 3,563 312 Updated Jul 13, 2026

A curated list of best cuda programming books

943 31 Updated May 19, 2026

Machine Learning Engineering Open Book

Python 18,463 1,180 Updated Jul 21, 2026

Module, Model, and Tensor Serialization/Deserialization

Python 318 53 Updated Jul 7, 2026

Benchmark suite for LLMs from Fireworks.ai

Python 111 38 Updated Jul 24, 2026

High Performance LLM Inference Operator Library

C++ 1,063 119 Updated Jul 24, 2026

📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉

Python 5,415 428 Updated Jun 23, 2026

SkyRL: A Modular Full-stack RL Library for LLMs

Python 2,092 389 Updated Jul 25, 2026

LLMPerf is a library for validating and benchmarking LLMs

Python 1,129 203 Updated Dec 9, 2024

Manages Unified Access to Generative AI Services built on Envoy Gateway

Go 1,865 316 Updated Jul 23, 2026
C++ 545 46 Updated Jul 14, 2026

Inference server benchmarking tool

Rust 165 33 Updated Jun 9, 2026

Fluid, elastic data abstraction and acceleration for BigData/AI applications in cloud. (Project under CNCF)

Go 1,956 1,271 Updated Jul 23, 2026

Using CRDs to manage GPU resources in Kubernetes.

Go 214 30 Updated Nov 21, 2022

Heterogeneous GPU Sharing on Kubernetes

Go 4,058 635 Updated Jul 24, 2026

AI on GKE is a collection of examples, best-practices, and prebuilt solutions to help build, deploy, and scale AI Platforms on Google Kubernetes Engine

Jupyter Notebook 329 243 Updated Jun 23, 2025

Cloud Native Benchmarking of Foundation Models

Python 46 20 Updated Jul 31, 2025

A fast GPU memory copy library based on NVIDIA GPUDirect RDMA technology

C 1,401 193 Updated Jul 14, 2026

LLM KV cache compression made easy

Python 1,146 160 Updated Jul 9, 2026

Open source AI coding agent. Designed for large projects and real world tasks.

Go 15,542 1,166 Updated Oct 3, 2025

A CLI inspector for the Model Context Protocol

JavaScript 441 41 Updated Jun 8, 2026

Serving multiple LoRA finetuned LLM as one

Python 1,168 65 Updated May 8, 2024
Next