Highlights
- Pro
Stars
Automated auditing pipeline for LLM and agent benchmarks — surfaces task ambiguity, environment conflicts, and evaluation bugs.
C++-based high-performance parallel environment execution engine (vectorized env) for general RL environments.
Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your claude code/codex/gemini agent will be an AI research agent with full horsepowe…
ARIS ⚔️ (Auto-Research-In-Sleep) — Lightweight Markdown-only skills for autonomous ML research: cross-model review loops, idea discovery, and experiment automation. No framework, no lock-in — works…
A holistic framework for advancing LLMs as data science agents
GRPO training code which scales to 32xH100s for long horizon terminal/coding tasks. Base agent is now the top Qwen3 agent on Stanford's TerminalBench leaderboard.
One local control plane for every AI agent: route across models, fuse new capabilities, orchestrate tools, and stay fully in control.
Claude Code. Any Model. The most powerful AI coding agent now speaks every language.
🔥 Comprehensive survey on Context Engineering: from prompt engineering to production-grade AI systems. hundreds of papers, frameworks, and implementation guides for LLMs and AI agents.
Examples of my Claude Code infrastructure with skill auto-activation, hooks, and agents
An interface library for RL post training with environments.
Efficient Triton Kernels for LLM Training
slime is an LLM post-training framework for RL Scaling.
Awesome curated collection of images and prompts generated by gemini-2.5-flash-image (aka Nano Banana) state-of-the-art image generation and editing model. Explore AI generated visuals created with…
A curated list of awesome Deep Reinforcement Learning resources.
[NeurIPS 2025 Spotlight] Reasoning Environments for Reinforcement Learning with Verifiable Rewards
OpenSpiel is a collection of environments and algorithms for research in general reinforcement learning and search/planning in games.
This repository contains a curated collection of 300+ case studies from over 80 companies, detailing practical applications and insights into machine learning (ML) system design. The contents are o…
Archon provides a modular framework for combining different inference-time techniques and LMs with just a JSON config file.
Super-Efficient RLHF Training of LLMs with Parameter Reallocation
A framework for few-shot evaluation of language models.
[NeurIPS 2024] SimPO: Simple Preference Optimization with a Reference-Free Reward
Together Mixture-Of-Agents (MoA) – 65.1% on AlpacaEval with OSS models
A quick guide (especially) for trending instruction finetuning datasets
[ACL 2024] Raccoon: Prompt Extraction Benchmark of LLM-Integrated Applications