Stars
Elemental Diagnosis of Generalist Mobile Manipulation Policies
AgentENV (AENV) is a distributed platform for running agent environments at scale.
A benchmark for evaluating LLMs on open-ended CS problems. Exploring the Next Frontier of Computer Science.
Measuring and evolving with the frontier of agent work
FORTE (Full-cycle Office Real-world Task Evaluation) is a general agent benchmark for evaluating AI agents on daily office productivity across 15 corporate professions.
Solve puzzles. Improve your pytorch.
Research artifacts from Recursive's automated AI research system
MiMo Code: Where Models and Agents Co-Evolve
CUA-Gym-Hub: mock web apps as reproducible RL training environments for computer-use agents
Scalable pipeline for synthesizing verifiable RLVR training data for computer-use agents
Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining (ICML2026)
[arXiv 2026] Learning from Rare Success and Rich Feedback via Reflection-Enhanced Self-Distillation
Can Language Models Rebuild Programs From Scratch?
1st Multilingual Benchmark for Repository-Level E2E Microservice Generation
Lightweight coding agent that runs in your terminal
Ultra-lightweight, open-source, self-hosted personal AI agent framework in Python with WebUI, tools, memory, MCP, multi-agent workflows, automation, and chat apps
Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
Official code for "SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization"
Official skills for the GLM family of models.
Code for the paper: Sketch2Code: Evaluating Vision-Language Models for Interactive Web Design Prototyping
0 - 1 learn OpenClaw: sections to build an claw-AI agent from scratch
Bash is all you need - A nano claude code–like 「agent harness」, built from 0 to 1
AI agents running research on single-GPU nanochat training automatically