Stars
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing,…
An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.
Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
A synthetic tabular and relational data generation framework
fmchisel: Efficient Compression and Training Algorithms for Foundation Models
AdalFlow: The library to build & auto-optimize LLM applications.
GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
Bash is all you need - A nano claude code–like 「agent harness」, built from 0 to 1
An extremely fast Python package and project manager, written in Rust.
slime is an LLM post-training framework for RL Scaling.
Official PyTorch implementation for "Large Language Diffusion Models"
Official implementation of "Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding"
3x Faster Inference; Unofficial implementation of EAGLE Speculative Decoding
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
Robust Speech Recognition via Large-Scale Weak Supervision
Train transformer language models with reinforcement learning.
LightLLM is a Python-based LLM (Large Language Model) inference and serving framework, notable for its lightweight design, easy scalability, and high-speed performance.
Production-tested AI infrastructure tools for efficient AGI development and community-driven innovation
🐳 Efficient Triton implementations for "Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention"
verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
High-performance inference framework for large language models, focusing on efficiency, flexibility, and availability.
A fast communication-overlapping library for tensor/expert parallelism on GPUs.
Collection of best practices, reference architectures, model training examples and utilities to train large models on AWS.
Best practices & guides on how to write distributed pytorch training code
Puzzles for learning Triton