Stars
SGLang is a high-performance serving framework for large language models and multimodal models.
DeepGEMM: clean and efficient BLAS kernel library on GPU
Entrpi/ds4, a Blackwell CUDA perf fork of antirez/ds4 on NVIDIA DGX Spark: one-command install, ~3x upstream prefill, ~1.5x decode, DSpark, and full continuous batch support
Reference implementation and examples of the CuTe Layout representation and algebra.
100M tokens. Infinite compute. Lowest val loss wins.
CODA: Rewriting Transformer Blocks as GEMM-Epilogue Programs
Long Context Pre-Training with Lighthouse Attention
Cuda kernels for leveraging LLM sparsity to improve throughput and decrease the memory requirements during inference and training.
Muon is an optimizer for hidden layers in neural networks
SmoothE: Differentiable E-Graph Extraction (ASPLOS'25 Best Paper)
TokenSpeed is a speed-of-light LLM inference engine.
cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.
Node0: A collaborative event powered by Protocol Learning, our decentralized approach to AI development
Optimized GPU compiler for LLM inference. Choose from a list of optimized recipes or optimize your own model via kernel fusion, autotuning, and advanced scheduling. Run benchmarks across different …
[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.
A beautiful, simple, clean, and responsive Jekyll theme for academics
just-every / code
Forked from openai/codexEvery Code - push frontier AI to it limits. A fork of the Codex CLI with validation, automation, browser integration, multi-agents, theming, and much more. Orchestrate agents from OpenAI, Claude, G…
FlashKDA: high-performance Kimi Delta Attention kernels
how few training tokens can you use to reach a target validation loss?
Accelerating MoE with IO and Tile-aware Optimizations