Starred repositories
Executable, measurable, and reproducible AI4AI toward recursive self-improvement. Home of OpenMLE and Frontis-MA1.
AI agents running research on single-GPU nanochat training automatically
KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+ more DSLs)
LLM-as-a-Verifier is a general-purpose framework that provides fine-grained feedback for any agent without requiring additional training. It achieves SOTA performance across coding, robotics, and m…
Programmable datacenter-scale infrastructure for Agents.
分享AI Infra知识&代码练习:PyTorch、vLLM/SGLang、slime/vime框架入门⚡️、性能加速🚀、大模型基础🧠、AI软硬件🔧等
Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. Tensor…
Achieve state of the art inference performance with modern accelerators on Kubernetes
[MLsys2026]: RAG on Everything with LEANN. Enjoy 97% storage savings while running a fast, accurate, and 100% private RAG application on your personal device.
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
Distribute and run AI workloads on Kubernetes magically in Python, like PyTorch for ML infra.
RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.
The agent that grows with you
Training library for Megatron-based models with bidirectional Hugging Face conversion capability
Ongoing research training transformer models at scale
JaxPP is a library for JAX that enables flexible MPMD pipeline parallelism for large-scale LLM training
Composable transformations of Python+NumPy programs: differentiate, vectorize, JIT to GPU/TPU, and more
NVIDIA / maxtext-jaxpp
Forked from AI-Hypercomputer/maxtextShowcase JaxPP with MaxText
A machine learning compiler for GPUs, CPUs, and ML accelerators
Orbax provides common checkpointing and persistence utilities for JAX users
A profiling and performance analysis tool for machine learning
TPU inference for vLLM, with unified JAX and PyTorch support.