Highlights
Starred repositories
Lean 4 programming language and theorem prover
Tile primitives for speedy kernels
A self-improving RLM agent for coding workflows and long-running autonomous tasks.
Synthetic data generation, post-training, and E2B benchmark evaluation infrastructure.
Incredibly fast JavaScript runtime, bundler, test runner, and package manager – all in one
Every past session, subagent, and workflow -- queryable by your agent, browsable by you
Mixture-of-experts (MoE) training megakernel for NVL72s
Wiki built using Karpathy method containing information about TPU performance optimizations and hooking it up to autoresearch optimization engine
Fast and Memory-Efficient Exact Attention for Large Headdim, 1.5x~6x speedup over PyTorch SDPA.
Lightweight loop engineering state kernel for long-running AI agent teams. Agent-loop agnostic across Codex, Claude Code, and other coding agents, with durable goals, quota-aware auto-wake, executa…
Agent-native evolutionary optimization. Say the goal in english, evolve code toward a measured target with a population of coding agents.
Puzzles for learning Triton, play it with minimal environment configuration!
🍎 One kernel a day keeps high latency away. A hands-on CUDA learning path featuring a rich collection of kernels, from the basics to peak performance, seamlessly integrated as PyTorch C++ extensions.
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention
Optimized local serving engine for Kimi-Linear-48B: INT4 quantizer, fused decode kernels for a measured 3.18x, and an OpenAI-compatible server. Ships with k3, a bridge that detects the client per r…
Native Blackwell (sm_100) tcgen05 training backward for the gated-linear-recurrence family (GDN-2/GLA/KDA/SSD), plus a contract-grade verifier that falsifies published GPU kernels. Six open Mamba-3…
Megatron Lite support for moonshotai/Kimi-K3 (KDA + gated MLA hybrid attention, LatentMoE, MXFP4 weights) — external model integration example
MoonEP: A Perfectly Balanced Expert Parallelism Library via Dynamic Redundant Experts
AgentENV (AENV) is a distributed platform for running agent environments at scale.
Itssshikhar / Flash-Flash-KDA
Forked from MoonshotAI/FlashKDAFlash-Flash KDA: H100-optimized Flash Kimi Delta Attention kernels
An evaluation benchmark for undergraduate competition math in Lean4, Isabelle, Coq, and natural language.
Welcome to KernelBench-Verified. This repository provides a robust, realistic evaluation framework for assessing custom CUDA kernels generated by Large Language Models (LLMs).
欢迎来到电子书下载宝库,一个汇聚了各类电子书下载链接的地方。无论你是喜欢阅读经典文学、经管励志、终身学习、职场创业、技术手册还是其他类型的书籍,这里都能满足你的需求。 该库涵盖了帆书app(原樊登读书)、微信读书、京东读书、喜马拉雅等读书app的大部分电子书。
Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM
Same model, different wrapper: a from-scratch benchmark comparing coding-agent harnesses (codex, pi, opencode, cursor, devin) and open models on correctness, speed, and token cost