Starred repositories
👻 Ghostty is a fast, feature-rich, and cross-platform terminal emulator that uses platform-native UI and GPU acceleration.
Meta-Framework of Spatiotemporal Composability
DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
The fastest, litest AI Gateway. Rust core with Python SDK. Call 100+ LLM APIs in OpenAI (or native) format with cost tracking, guardrails, load balancing, and logging [Bedrock, Azure, OpenAI, Anthr…
TokenSpeed is a speed-of-light LLM inference engine.
Use Codex from Claude Code to review code or delegate tasks.
A Claude Code plugin that shows what's happening - context usage, active tools, running agents, and todo progress
Garry's Opinionated OpenClaw/Hermes Agent Brain
Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.
Tile-Based Runtime for Ultra-Low-Latency LLM Inference
Open-source book with Modern CUDA Learn Notes for Beginners, includes FP16/BF16, FP8, HGEMM, FlashAttention, CuTe, etc.
slime is an LLM post-training framework for RL Scaling.
Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels
The source of LMSYS website and blogs
A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downs…
A fast communication-overlapping library for tensor/expert parallelism on GPUs.
Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, DeepSeek-V4, MiniMax-H3, Gemma 4, FLUX and more.
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. Tensor…
Mirage Persistent Kernel: Compiling LLMs into a MegaKernel
DuckLake is an integrated data lake and catalog format
Analyze computation-communication overlap in V3/R1.
Documented system prompts from Anthropic - Claude Fable 5.1, Opus 5.5, Claude Design, Claude Code. OpenAI - ChatGPT GPT-6-Astra, Codex. Google - Gemini 3.8 Flash, 3.1 Pro, Antigravity. xAI - Grok, …
Production-grade client-side tracing, profiling, and analysis for complex software systems.
A high-throughput and memory-efficient inference and serving engine for LLMs
My learning notes for ML SYS.