Skip to content
View TobiasLee's full-sized avatar
🎯
Focusing
🎯
Focusing

Organizations

@lancopku

Block or report TobiasLee

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Elemental Diagnosis of Generalist Mobile Manipulation Policies

HTML 129 5 Updated Jul 28, 2026

Agentic RL Training at Scale

Python 1,943 406 Updated Aug 18, 2026

AgentENV (AENV) is a distributed platform for running agent environments at scale.

Rust 3,216 270 Updated Aug 18, 2026

A benchmark for evaluating LLMs on open-ended CS problems. Exploring the Next Frontier of Computer Science.

C++ 295 44 Updated Aug 13, 2026

LongCLI-Bench's official repository

Python 44 3 Updated Aug 14, 2026

Measuring and evolving with the frontier of agent work

Python 507 383 Updated Aug 17, 2026

FORTE (Full-cycle Office Real-world Task Evaluation) is a general agent benchmark for evaluating AI agents on daily office productivity across 15 corporate professions.

Python 19 1 Updated Jun 30, 2026

Solve puzzles. Improve your pytorch.

Jupyter Notebook 4,288 392 Updated Jul 15, 2024

Advancing Open-source World Models

Python 4,371 396 Updated Jul 9, 2026

Research artifacts from Recursive's automated AI research system

Python 209 19 Updated Jun 11, 2026

MiMo Code: Where Models and Agents Co-Evolve

TypeScript 12,788 1,309 Updated Aug 18, 2026

CUA-Gym-Hub: mock web apps as reproducible RL training environments for computer-use agents

JavaScript 71 10 Updated Jul 27, 2026

Scalable pipeline for synthesizing verifiable RLVR training data for computer-use agents

Python 186 17 Updated Aug 13, 2026

Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining (ICML2026)

JavaScript 38 Updated Jun 29, 2026

[arXiv 2026] Learning from Rare Success and Rich Feedback via Reflection-Enhanced Self-Distillation

Python 18 1 Updated May 21, 2026

Can Language Models Rebuild Programs From Scratch?

Python 898 63 Updated Jul 26, 2026

1st Multilingual Benchmark for Repository-Level E2E Microservice Generation

Python 100 9 Updated Apr 20, 2026

Lightweight coding agent that runs in your terminal

Rust 106,660 16,199 Updated Aug 18, 2026

Ultra-lightweight, open-source, self-hosted personal AI agent framework in Python with WebUI, tools, memory, MCP, multi-agent workflows, automation, and chat apps

Python 47,144 8,327 Updated Aug 18, 2026

Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding

Python 365 4 Updated Aug 5, 2026

Official code for "SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization"

Python 364 17 Updated Aug 12, 2026

Official skills for the GLM family of models.

Python 461 40 Updated Apr 15, 2026

Code for the paper: Sketch2Code: Evaluating Vision-Language Models for Interactive Web Design Prototyping

Python 42 8 Updated Oct 29, 2024

0 - 1 learn OpenClaw: sections to build an claw-AI agent from scratch

Python 3,270 381 Updated Jun 30, 2026

Bash is all you need - A nano claude code–like 「agent harness」, built from 0 to 1

Python 74,573 12,063 Updated Aug 18, 2026

AI agents running research on single-GPU nanochat training automatically

Python 94,078 13,323 Updated Mar 26, 2026
Next