Stars
High-performance GEMM kernel examples with FlyDSL on AMD GPUs.
Multi-LLM peer review for code decisions. Bring your own CLI; Chorus convenes 2-4 other LLMs to review the work before you ship.
A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations
Agents, and RL environment, for optimizing GPU kernels on AMD ROCm using LLM agents. Benchmarks LLM serving workloads end-to-end, profiles bottleneck kernels, optimizes them via Claude Code or Code…
A curriculum for learning about gpu performance engineering, from scratch to what the frontier AI labs do
AGENTS.md/CLAUDE.md generator and ExecPlan harness for Codex, Claude Code, Cursor, Antigravity, OpenCode, Pi and other coding agents. Use api keys or authenticate with Codex or Claude Code
Machine Learning Engineering Open Book
Automated bottleneck detection and solution orchestration
A simple yet powerful tool to turn traditional container/OS images into unprivileged sandboxes.
A lightweight, local-first, and free experiment tracking library from Hugging Face 🤗
A Python-embedded DSL that makes it easy to write fast, scalable ML kernels with minimal boilerplate.
Split-screen video comparison tool using FFmpeg and SDL2
Autonomous coding agent as an SDK, IDE extension, or CLI assistant.
A framework for few-shot evaluation of language models.
An interactive web-based tool for exploring intermediate representations of PyTorch and Triton models
JAX library for training sub-4B foundation models for edge
Flax (Jax) implementation of DeepSeek-R1-Distill-Qwen-1.5B with weights ported from Hugging Face.
Repository to host ROCm Developer Hub Notebook Tutorials
NUS CS5242 Neural Networks and Deep Learning, Xavier Bresson, 2025
A minimal, single-file implementation of the Mamba-2 model in JAX.
ROCm / triton
Forked from triton-lang/tritonDevelopment repository for the Triton language and compiler
Framework to reduce autotune overhead to zero for well known deployments.
Deep learning for dummies. All the practical details and useful utilities that go into working with real models.