Stars
Design of noise, interpolation schedules, and sources in generative dynamics, with enhanced numerical performance
The codebase of our paper "Improving the Training of Rectified Flows", NeurIPS 2024
State-of-the-Art Embeddings, Retrieval, and Reranking
Emergent Hierarchical Reasoning in LLMs/VLMs through Reinforcement Learning [ICLR26]
Code for Fine-grained Uncertainty Quantification for LLMs from Semantic Similarities (NeurIPS'24)
a neuroscience based model simulating basal ganglia for mouse maze pathfinding with q-learning
Minimal reproduction of DeepSeek R1-Zero
An Open-source RL System from ByteDance Seed and Tsinghua AIR
The Entropy Mechanism of Reinforcement Learning for Large Language Model Reasoning.
Train transformer language models with reinforcement learning.
[NeurIPS 2025] TTRL: Test-Time Reinforcement Learning
Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends
verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
PyTorch implementation of Advantage Actor Critic (A2C), Proximal Policy Optimization (PPO), Scalable trust-region method for deep reinforcement learning using Kronecker-factored approximation (ACKT…
Code for NeurIPS'24 paper 'Grokked Transformers are Implicit Reasoners: A Mechanistic Journey to the Edge of Generalization'
Code for the paper "The Impact of Positional Encoding on Length Generalization in Transformers", NeurIPS 2023
An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Cliff walking reinforcement learning example, with a variety of RL algorithms
The simplest, fastest repository for training/finetuning medium-sized GPTs.