Starred repositories
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing,…
Recipes and resources for building, deploying, and fine-tuning generative AI with Fireworks.
Model compression toolkit engineered for enhanced usability, comprehensiveness, and efficiency.
DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms
DeepEP: an efficient expert-parallel communication library
mKernel: fast multi-node, multi-GPU fused kernels
Tile primitives for speedy kernels
A curated list of best cuda programming books
Machine Learning Engineering Open Book
Module, Model, and Tensor Serialization/Deserialization
High Performance LLM Inference Operator Library
📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉
SkyRL: A Modular Full-stack RL Library for LLMs
LLMPerf is a library for validating and benchmarking LLMs
Manages Unified Access to Generative AI Services built on Envoy Gateway
Inference server benchmarking tool
Fluid, elastic data abstraction and acceleration for BigData/AI applications in cloud. (Project under CNCF)
Using CRDs to manage GPU resources in Kubernetes.
AI on GKE is a collection of examples, best-practices, and prebuilt solutions to help build, deploy, and scale AI Platforms on Google Kubernetes Engine
Cloud Native Benchmarking of Foundation Models
A fast GPU memory copy library based on NVIDIA GPUDirect RDMA technology
Open source AI coding agent. Designed for large projects and real world tasks.
A CLI inspector for the Model Context Protocol