Stars
- All languages
- AGS Script
- C
- C#
- C++
- CMake
- CSS
- CoffeeScript
- Cuda
- Cython
- Dockerfile
- GLSL
- Go
- Graphviz (DOT)
- HTML
- Java
- JavaScript
- Julia
- Jupyter Notebook
- Kotlin
- LLVM
- Lua
- MATLAB
- MDX
- MLIR
- Markdown
- Objective-C
- Objective-C++
- OpenSCAD
- PHP
- PLpgSQL
- Perl
- PowerShell
- Protocol Buffer
- Pug
- PureBasic
- Python
- R
- Roff
- Ruby
- Rust
- Scala
- Shell
- Swift
- SystemVerilog
- TeX
- TypeScript
- Vim Script
- Vue
FlashKDA: high-performance Kimi Delta Attention kernels
MoonEP: A Perfectly Balanced Expert Parallelism Library via Dynamic Redundant Experts
OpenClaw-RL: Train any agent simply by talking
An LLM post-training framework with vLLM for RL Scaling
slime is an LLM post-training framework for RL Scaling.
A safetensors extension to efficiently store sparse quantized tensors on disk
meituan-longcat / mscclpp
Forked from microsoft/mscclppMSCCL++: A GPU-driven communication stack for scalable AI applications
The Agent Harness for AI-Human Collaboration, inspired by the AI-DLC (AI-Driven Development Lifecycle)
tile-ai / tilescale
Forked from tile-ai/tilelangTile-based language built for AI computation across all scales
High-performance LLM operator library built on TileLang.
Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels
Tile-Based Runtime for Ultra-Low-Latency LLM Inference
aws-neuron / torchtitan-neuron
Forked from pytorch/torchtitanA PyTorch native platform for training generative AI models
Source Han Sans | 思源黑体 | 思源黑體 | 思源黑體 香港 | 源ノ角ゴシック | 본고딕
The official repo for STCast (CVPR2026 Highlight).
🚀 Efficient implementations for emerging model architectures
abrooks98 / nvshmem
Forked from NVIDIA/nvshmemNVIDIA NVSHMEM is a parallel programming interface for NVIDIA GPUs based on OpenSHMEM. NVSHMEM can significantly reduce multi-process communication and coordination overheads by allowing programmer…
An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
rauteric / DeepEP
Forked from deepseek-ai/DeepEPDeepEP: an efficient expert-parallel communication library
NVIDIA NVSHMEM is a parallel programming interface for NVIDIA GPUs based on OpenSHMEM. NVSHMEM can significantly reduce multi-process communication and coordination overheads by allowing programmer…
amazon-contributing / upstream-to-nccl
Forked from NVIDIA/ncclOptimized primitives for collective multi-GPU communication
MSCCL++: A GPU-driven communication stack for scalable AI applications