Stars
A minimal implementation of DeepMind's Genie world model
GRPO training code which scales to 32xH100s for long horizon terminal/coding tasks. Base agent is now the top Qwen3 agent on Stanford's TerminalBench leaderboard.
The simplest, fastest repository for training/finetuning medium-sized GPTs.
Composable transformations of Python+NumPy programs: differentiate, vectorize, JIT to GPU/TPU, and more
An educational resource to help anyone learn deep reinforcement learning.
Machine learning on FPGAs using HLS