Stars
CLI proxy that reduces LLM token consumption by 60-90% on common dev commands. Single Rust binary, zero dependencies
Lightweight coding agent that runs in your terminal
Tile-Based Runtime for Ultra-Low-Latency LLM Inference
FlyDSL is the Python front‑end of the project: a Flexible Layout Python DSL for expressing tiling, partitioning, data movement, and kernel structure at a high level.
An agentic skills framework & software development methodology that works.
high-performance linear attention kernel library built on TileLang
KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+ more DSLs)
FlashKDA: high-performance Kimi Delta Attention kernels
High-performance GEMM kernel examples with FlyDSL on AMD GPUs.
A cross-platform desktop All-in-One assistant for Claude Code, Codex, OpenCode, OpenClaw, Grok Build & Hermes Agent. Only official website: ccswitch.io
Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflo…
Autonomous GPU Kernel Generation & Optimization via Deep Agents
CUDA kernels for linear attention variants, written in CuTe DSL and CUTLASS C++.
🚀 Efficient implementations for emerging model architectures
Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.
A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.
RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.
Mirage Persistent Kernel: Compiling LLMs into a MegaKernel
CUDA Python: Performance meets Productivity
An Emacs framework for the stubborn martian hacker
DeepEP: an efficient expert-parallel communication library
A Datacenter Scale Distributed Inference Serving Framework
[DEPRECATED] Moved to ROCm/rocm-libraries repo. NOTE: develop branch is maintained as a read-only mirror
DeepGEMM: clean and efficient BLAS kernel library on GPU