Stars
kernelboard is the webapp for https://www.gpumode.com
KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+ more DSLs)
Autonomous GPU Kernel Generation & Optimization via Deep Agents
Local UI to run and train LLMs and diffusion models, including Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, DeepSeek-V4, FLUX and more.
A high-throughput and memory-efficient inference and serving engine for LLMs
Analyze videos using LLMs, Computer Vision and Automatic Speech Recognition
Run PyTorch LLMs locally on servers, desktop and mobile
A framework for few-shot evaluation of language models.
On-device AI across mobile, embedded and edge for PyTorch
Tensors and Dynamic neural networks in Python with strong GPU acceleration