Highlights
Lists (10)
Sort Name ascending (A-Z)
Stars
Annotate and review coding agent plans and code diffs visually, share with your team, send feedback to agents with one click.
agent multiplexer that lives in your terminal.
llm-d Router: The intelligent entry point for inference requests
Let your coding agents run wild in parallel: a fully isolated dev environment per git worktree, ports, proxy, and databases
🚀 A simple way to launch, train, and use PyTorch models on almost any device and distributed configuration, automatic mixed precision (including fp8), and easy-to-configure FSDP and DeepSpeed support
Post-training with Tinker
A course of learning LLM inference serving on Apple Silicon for systems engineers: build a tiny vLLM + Qwen.
Measure tokens/sec of any LLM behind an OpenAI-compatible API.
Textbook on reinforcement learning from human feedback
⌥ AI Coding agent for the terminal — hash-anchored edits, optimized tool harness, LSP, Python, browser, subagents, and more
slime is an LLM post-training framework for RL Scaling.
A PyTorch native platform for training generative AI models
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
Achieve state of the art inference performance with modern accelerators on Kubernetes
SIA is a Self Improving AI framework to autonomously improve the performance of any AI system (Model / Agent) on a benchmark task.
A curated list of best cuda programming books
Programmable chat templates for LLM training and inference.
Experimental implementation of DeepSeek v4 flaash in llama.cpp
A Python-embedded DSL that makes it easy to write fast, scalable ML kernels with minimal boilerplate.
Autoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton kernels.
Create stunning demos for free. Open-source, no subscriptions, no watermarks, and free for commercial use. An alternative to Screen Studio.
ripgrep recursively searches directories for a regex pattern while respecting your gitignore