Lists (1)
Sort Name ascending (A-Z)
Stars
Frequently updated list of dLLM (Diffusion Large Language Models) papers, models, and other resources
DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms
A compiler, optimizer and executor for financial expressions and factors
TurboQuant: Near-optimal KV cache quantization for LLM inference (3-bit keys, 2-bit values) with Triton kernels + vLLM integration
A kernel library written in tilelang
An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.
PTX ISA 9.1 documentation converted to searchable markdown. Includes Claude Code skill for CUDA development.
IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse
A lightweight inference engine supporting speculative speculative decoding (SSD).
SGLang Omni: High-Performance Multi-Stage Pipeline Framework for Omni Models
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, …
Autonomous GPU Kernel Generation & Optimization via Deep Agents
CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation
DFlash: Block Diffusion for Flash Speculative Decoding
Nsight Python is a Python kernel profiling interface based on NVIDIA Nsight Tools
Our first fully AI generated deep learning system
Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
[ICLR 2025] DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
Accelerating Large-Scale Reasoning Model Inference with Sparse Self-Speculative Decoding
Accelerating MoE with IO and Tile-aware Optimizations
SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse–Linear Attention
Tile-Based Runtime for Ultra-Low-Latency LLM Inference