Lists (3)
Sort Name ascending (A-Z)
Stars
SilentPatch for GTA III, Vice City, and San Andreas
Mixture-of-experts (MoE) training megakernel for NVL72s
MSLK (Meta Superintelligence Labs Kernels) is a collection of PyTorch GPU operator libraries that are designed and optimized for GenAI training and inference, such as FP8 row-wise quantization and …
Arm Performix: https://developer.arm.com/servers-and-cloud-computing/arm-performix
Reference implementation and examples of the CuTe Layout representation and algebra.
MoonEP: A Perfectly Balanced Expert Parallelism Library via Dynamic Redundant Experts
Itssshikhar / Flash-Flash-KDA
Forked from MoonshotAI/FlashKDAFlash-Flash KDA: H100-optimized Flash Kimi Delta Attention kernels
A feed-forward 3D foundation model for reconstructing scenes from streaming data
SpaceXAI's coding agent harness and TUI. Fullscreen, mouse interactive, extensible.
Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦
A web engine for pencil-style rendering of 3D scenes to SVG — exact silhouettes, hidden-line ghosting, hatching. Great for abstract and technical scenes.
Cross-platform instrumentation and introspection library written in C
Cross-Platform HW accelerated CRC32c and CRC32 with fallback to efficient SW implementations. C interface with language bindings for each of our SDKs
Standalone C++/GGML runtime for ThinkSound text->sound-effect generation
A comprehensive collection of IQA papers
slime is an LLM post-training framework for RL Scaling.
High-performance code intelligence MCP server. Indexes codebases into a persistent knowledge graph — average repo in milliseconds. 158 languages, sub-ms queries, 99% fewer tokens. Single static bin…
Low-latency Rust thread pool with parallel iterators
C++ implementation of a fast hash map and hash set using robin hood hashing
Real-time 3D full-body reconstruction from a single camera, Multiperson BVH output, Pure C++ runtime, ONNX + ggml, 70-joint skeleton with hands.
FlashKDA: high-performance Kimi Delta Attention kernels
DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
TokenSpeed is a speed-of-light LLM inference engine.