Stars
An agentic system that auto-optimizes LLM workloads on AMD GPUs.
Toolkit for launching and observing MaxText training on Slurm-managed GPU clusters
A high-performance acceleration library dedicated to large-scale model training on AMD GPUs
A high-performance distributed file system designed to address the challenges of AI training and inference workloads.
UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., GPU-driven)
A flexible and high-performance training framework designed for large-scale foundation model training on AMD GPUs
Composable transformations of Python+NumPy programs: differentiate, vectorize, JIT to GPU/TPU, and more
Master the command line, in one page
Apache Beam is a unified programming model for Batch and Streaming data processing.