Stars
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing,…
TokenSpeed is a speed-of-light LLM inference engine.
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
Run Kubernetes on MySQL, Postgres, sqlite, not etcd.
You like pytorch? You like micrograd? You love tinygrad! ❤️
🚀 Pytorch Distributed native training library for LLMs/VLMs with OOTB Hugging Face support
A curated collection of papers and resources on On-Policy Distillation for Large Language Models.
Production-speed compact Dynamic Memory Sparsification (DMS) for KV cache compression
Markdown Input & Text for React Native
A flexible distributed key-value database that is optimized for caching and other realtime workloads.
Integrates tracing and logging with egui for event collection/visualization
A Python-embedded DSL that makes it easy to write fast, scalable ML kernels with minimal boilerplate.
slime is an LLM post-training framework for RL Scaling.
A Datacenter Scale Distributed Inference Serving Framework
Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究平台…
A Git-compatible VCS that is both simple and powerful
Scalable toolkit for efficient model reinforcement
Training library for Megatron-based models with bidirectional Hugging Face conversion capability
Production-tested AI infrastructure tools for efficient AGI development and community-driven innovation
A simple program that emulates the detach feature of screen
From the creators of Unistyles: The fastest Tailwind bindings for React Native
React Native module for real-time screenshot detection on Android and iOS
Official TypeScript API client library for turbopuffer