Skip to content
View andyxning's full-sized avatar
🎯
Focusing
🎯
Focusing
  • Beijing, China

Organizations

@nsqio @kubernetes

Block or report andyxning

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Starred repositories

Showing results

Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing,…

Rust 421 126 Updated Jul 28, 2026

Recipes and resources for training, building generative AI with Fireworks

Jupyter Notebook 187 52 Updated Jul 28, 2026

Model compression toolkit engineered for enhanced usability, comprehensiveness, and efficiency.

Python 1,491 170 Updated Jul 28, 2026

DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms

Python 6,804 633 Updated Jul 9, 2026

DeepEP: an efficient expert-parallel communication library

Cuda 9,913 1,348 Updated Jul 27, 2026

mKernel: fast multi-node, multi-GPU fused kernels

Cuda 257 24 Updated Jul 28, 2026

Tile primitives for speedy kernels

Cuda 3,570 314 Updated Jul 13, 2026

A curated list of best cuda programming books

941 31 Updated May 19, 2026

Machine Learning Engineering Open Book

Python 18,482 1,183 Updated Jul 28, 2026

Module, Model, and Tensor Serialization/Deserialization

Python 318 53 Updated Jul 7, 2026

Benchmark suite for LLMs from Fireworks.ai

Python 111 38 Updated Jul 27, 2026

High Performance LLM Inference Operator Library

C++ 1,071 121 Updated Jul 24, 2026

📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉

Python 5,425 428 Updated Jul 26, 2026

SkyRL: A Modular Full-stack RL Library for LLMs

Python 2,103 393 Updated Jul 28, 2026

LLMPerf is a library for validating and benchmarking LLMs

Python 1,129 203 Updated Dec 9, 2024

Manages Unified Access to Generative AI Services built on Envoy Gateway

Go 1,873 318 Updated Jul 28, 2026
C++ 546 46 Updated Jul 14, 2026

Inference server benchmarking tool

Rust 165 32 Updated Jun 9, 2026

Fluid, elastic data abstraction and acceleration for BigData/AI applications in cloud. (Project under CNCF)

Go 1,956 1,272 Updated Jul 28, 2026

Using CRDs to manage GPU resources in Kubernetes.

Go 214 30 Updated Nov 21, 2022

Heterogeneous GPU Sharing on Kubernetes

Go 4,098 664 Updated Jul 28, 2026

AI on GKE is a collection of examples, best-practices, and prebuilt solutions to help build, deploy, and scale AI Platforms on Google Kubernetes Engine

Jupyter Notebook 329 243 Updated Jun 23, 2025

Cloud Native Benchmarking of Foundation Models

Python 46 20 Updated Jul 31, 2025

A fast GPU memory copy library based on NVIDIA GPUDirect RDMA technology

C 1,400 194 Updated Jul 14, 2026

LLM KV cache compression made easy

Python 1,148 164 Updated Jul 27, 2026

Open source AI coding agent. Designed for large projects and real world tasks.

Go 15,542 1,166 Updated Oct 3, 2025

A CLI inspector for the Model Context Protocol

JavaScript 441 41 Updated Jun 8, 2026

Serving multiple LoRA finetuned LLM as one

Python 1,168 65 Updated May 8, 2024
Next