Lists (8)
Sort Name ascending (A-Z)
Stars
[KDD 2026] Implementation for the paper "RAPO: Expanding Exploration for LLM Agents via Retrieval-Augmented Policy Optimization"
SkillOpt-Lite and HarnessOpt: Optimize your skill or harness with one line of vibe
Learning Agentic Policy from Action Guidance
Evaluation harness for Apodex-1.0 on public deep-research benchmarks.
"QUEST: Training Frontier Deep Research Agents with Fully Synthetic Tasks"
Beyond Semantic Similarity: Rethinking Retrieval for Agentic Search via Direct Corpus Interaction
Implementation of SLIM, a framework of dynamics skill lifecycle management for agentic reinforcement learning
Official implementation for paper "Demystifying Reinforcement Learning for Long-Horizon Tool-Using Agents: A Comprehensive Recipe"
Mobile-Agent: The Powerful GUI Agent Family
UniScientist is designed to advance universal scientific research intelligence through a unified paradigm
RLAnything (ICML 2026) & AutoTool (ICML 2026), DemyAgent: Open-Source RL for LLMs and Agentic Scenarios
Dr. MAS is an end-to-end RL training framework for multi-agent LLM systems, supporting the co-training of multiple (heterogeneous) LLMs.
RLinf: Reinforcement Learning Infrastructure for Embodied and Agentic AI
Elevate your AI research writing, no more tedious polishing ✨
This repository contains the code and data for the paper "Chaining the Evidence: Robust Reinforcement Learning for Deep Search Agents with Citation-Aware Rubric Rewards".
DeepResearch Bench II (DRB2) is the follow-up to DeepResearch Bench, with a stronger focus on measuring the gap between deep research systems and human experts. It does so by decomposing expert-wri…
qqr is an RL training framework for open-ended agents.
We introduce BabyVision, a benchmark revealing the infancy of AI vision.
PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning
Resources and paper list for 'Scaling Environments for Agents'. This repository accompanies our survey on how environments contribute to agent intelligence.
Develop review and rebuttal agents for openreview website
Public quant internship repository, maintained by NUFT but available for everyone.
[ICLR 2026] InftyThink: Breaking the Length Limits of Long-Context Reasoning in Large Language Models
[ICLR'26] SketchThinker-R1: Towards Efficient Sketch-Style Reasoning in Large Multimodal Models
(ICLR'26 + Netflix) Rank-GRPO: Training LLM-based Conversational Recommender Systems with Reinforcement Learning
[ICLR 2026] VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications
Open source code for ICLR 2026 Paper: Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions
Pushing Test-Time Scaling Limits of Deep Search with Asymmetric Verification