Stars
Procedural data generators for verifiable reasoning, synthetic pretraining, post-training, evaluation, and RL.
A skill to stop your coding agent from burying the answer. ADHD-friendly output.
Implementation of the paper: Beyond Correlation: The impact of human uncertainty in measuring the effectiveness of automatic evaluation and LLM-as-a-judge
MORPH: PDE Foundation Models with Arbitrary Data Modality
This repo is meant to serve as a guide for Machine Learning/AI technical interviews.
Official Repo of "$X$-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding"
[CVPR 2026 Findings] V-GRPO: Online Reinforcement Learning for Denoising Generative Models Is Easier than You Think
Code, Data, and Model Outputs for the paper "This Treatment Works, Right? Evaluating LLM Sensitivity to Patient Question Framing in Medical QA"
Elucidating the Design Space of Flow Matching for Cellular Microscopy
[ICML 2026] Code for V1: Unifying Generation and Self-Verification for Parallel Reasoners.
[ICML 2025] Teaching Language Models to Critique via Reinforcement Learning
PyTorch building blocks for the OLMo ecosystem
Modeling, training, eval, and inference code for OLMo
Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours
Implement a ChatGPT-like LLM in PyTorch from scratch, step by step
A framework bridging cognitive science and LLM reasoning research to diagnose and improve how large language models reason, based on analysis of 192K model traces and 54 human think-aloud traces.
Fully Open-source Multimodal Language Models for Science Discovery
[EACL 2026] PaperSearchQA. Data generation pipeline for QA over scientific papers, suitable for RL training search agents