Highlights
- Pro
Lists (2)
Sort Name ascending (A-Z)
Starred repositories
Our library for RL environments + evals
[TMLR] A curated list of language modeling researches for code (and other software engineering activities), plus related datasets.
OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
UniRL is a Framework for Unified Multimodal Model Reinforcement Learning
【ICML2026 Spotlight】 T2PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning
"QUEST: Training Frontier Deep Research Agents with Fully Synthetic Tasks"
Official Implementation of VisualClaw: A Real-Time, Personalized Agent for the Physical World
MiMo Code: Where Models and Agents Co-Evolve
One Discrete Word for Visual Reasoning Overtakes Agentic and Latent Methods
Compiler Version Manager — A cross-platform C/C++ Compiler Version Manager for LLVM and GCC. Similar to rustup, nvm, or rbenv but for C/C++ toolchains.
Code for paper OpenWebRL: Online Multi-Turn Reinforcement Learning for Visual Web Agents
Agentic RL on Any Harness at Scale
Official repo of "MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents". It can be used to evaluate a GUI agent with a hierarchical manner across multiple platforms, includi…
[ICML 2026] Official implementation of Target-Oriented Pretraining Data Selection via Neuron-Activated Graph
The implementation for SIGIR 2026: Learning to Retrieve from Agent Trajectories.
The agent that grows with you
Official repository for Your Agent, Their Asset: A Real-World Safety Analysis of OpenClaw
AutoHarness: Automated Harness Engineering for AI Agents
Codebase for the work “Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?”
Co-evolving policy actors and experience extractors for efficient experience-driven agent RL
The batteries-included agent harness.
OpenShell is the safe, private runtime for autonomous AI agents.
[CVPR 2026] Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generation
Fully autonomous & self-evolving research from idea to paper. Chat an Idea. Get a Paper. 🦞