Skip to content
View WxxShirley's full-sized avatar
🤔
focus
🤔
focus

Block or report WxxShirley

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

[KDD 2026] Implementation for the paper "RAPO: Expanding Exploration for LLM Agents via Retrieval-Augmented Policy Optimization"

Python 16 Updated Jul 23, 2026
Python 1,353 156 Updated Aug 10, 2026

SkillOpt-Lite and HarnessOpt: Optimize your skill or harness with one line of vibe

Python 164 9 Updated Jul 22, 2026

Learning Agentic Policy from Action Guidance

Python 12 1 Updated May 13, 2026

Evaluation harness for Apodex-1.0 on public deep-research benchmarks.

Python 380 38 Updated Jun 8, 2026

"QUEST: Training Frontier Deep Research Agents with Fully Synthetic Tasks"

Python 244 22 Updated Aug 5, 2026

Beyond Semantic Similarity: Rethinking Retrieval for Agentic Search via Direct Corpus Interaction

Python 382 48 Updated Jun 1, 2026

Implementation of SLIM, a framework of dynamics skill lifecycle management for agentic reinforcement learning

Python 22 1 Updated May 12, 2026

Official implementation for paper "Demystifying Reinforcement Learning for Long-Horizon Tool-Using Agents: A Comprehensive Recipe"

Python 39 1 Updated May 12, 2026

Mobile-Agent: The Powerful GUI Agent Family

Python 9,064 908 Updated Jul 7, 2026

UniScientist is designed to advance universal scientific research intelligence through a unified paradigm

Python 169 13 Updated Mar 14, 2026

RLAnything (ICML 2026) & AutoTool (ICML 2026), DemyAgent: Open-Source RL for LLMs and Agentic Scenarios

Python 611 59 Updated Jun 12, 2026

Dr. MAS is an end-to-end RL training framework for multi-agent LLM systems, supporting the co-training of multiple (heterogeneous) LLMs.

Python 152 10 Updated Jul 27, 2026

RLinf: Reinforcement Learning Infrastructure for Embodied and Agentic AI

Python 4,501 649 Updated Aug 10, 2026

Elevate your AI research writing, no more tedious polishing ✨

32,777 2,424 Updated May 18, 2026

This repository contains the code and data for the paper "Chaining the Evidence: Robust Reinforcement Learning for Deep Search Agents with Citation-Aware Rubric Rewards".

Python 73 9 Updated Apr 8, 2026

DeepResearch Bench II (DRB2) is the follow-up to DeepResearch Bench, with a stronger focus on measuring the gap between deep research systems and human experts. It does so by decomposing expert-wri…

Python 79 3 Updated May 14, 2026

qqr is an RL training framework for open-ended agents.

Python 276 22 Updated Aug 5, 2026

We introduce BabyVision, a benchmark revealing the infancy of AI vision.

Python 237 10 Updated Jan 13, 2026

PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning

Python 338 14 Updated Feb 5, 2026

Resources and paper list for 'Scaling Environments for Agents'. This repository accompanies our survey on how environments contribute to agent intelligence.

72 4 Updated Jan 28, 2026

Develop review and rebuttal agents for openreview website

Python 7 Updated Dec 15, 2025

Public quant internship repository, maintained by NUFT but available for everyone.

OCaml 2,434 152 Updated Jul 30, 2026

[ICLR 2026] InftyThink: Breaking the Length Limits of Long-Context Reasoning in Large Language Models

Python 57 1 Updated May 5, 2026

[ICLR'26] SketchThinker-R1: Towards Efficient Sketch-Style Reasoning in Large Multimodal Models

Python 17 2 Updated Mar 26, 2026

(ICLR'26 + Netflix) Rank-GRPO: Training LLM-based Conversational Recommender Systems with Reinforcement Learning

Python 53 7 Updated May 23, 2026

[ICLR 2026] VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications

Python 163 17 Updated Feb 22, 2026

Open source code for ICLR 2026 Paper: Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions

Python 422 59 Updated May 21, 2026

Pushing Test-Time Scaling Limits of Deep Search with Asymmetric Verification

Python 22 1 Updated Oct 8, 2025
Next