- Shanghai, China
Stars
The fastest, litest AI Gateway. Rust core with Python SDK. Call 100+ LLM APIs in OpenAI (or native) format with cost tracking, guardrails, load balancing, and logging [Bedrock, Azure, OpenAI, Anthr…
Qwen-CUA: Native Computer Use for (Almost) Everything — a screenshot-driven agent that operates computers with keyboard and mouse, jointly developed by the Qwen Team and XLang Lab.
WebWorld is a large-scale web world model that helps train web agents in a simulated browser, avoiding the latency and safety issues of the real web.
BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions [TMLR2025]
WorkArena: How Capable are Web Agents at Solving Common Knowledge Work Tasks?
Fully automatic censorship removal for language models
Implementation of Reinforcement Learning Algorithms (From Reinforcement Learning An Introduction By Sutton & Barto)
Implementation of all RL algorithms in a simpler way
Test that your Claude Code skills, MCP servers, and CLIs actually work when an agent uses them — sandboxed YAML suites, activation checks, A/B experiments, CI gates.
CLI for common Playwright actions. Record and generate Playwright code, inspect selectors and take screenshots.
Lightweight coding agent that runs in your terminal
slime is an LLM post-training framework for RL Scaling.
Optimize prompts, code, and more with AI-powered Reflective Optimization
Official repository for SaaS-Bench: realistic, locally deployable SaaS workflows for GUI agent evaluation.
Give your AI agent access to your live Chrome session — works out of the box, connects to tabs you already have open
Official AHE code — Agentic Harness Engineering: observability-driven automatic evolution of coding-agent harnesses (concurrent w/ meta-harness). NexAU-AHE reaches 84.7% ± 2.1 pass@1 on Terminal-Be…
A library of HTML slide templates designed so any coding agent can pick the right one and produce a beautiful deck on the user's behalf, automatically.
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
Heuristic Learning Blog Post
RLAnything (ICML 2026) & AutoTool (ICML 2026), DemyAgent: Open-Source RL for LLMs and Agentic Scenarios
My learning notes for ML SYS.
Convert AI coding agent sessions (Claude Code, Cursor, Codex, Gemini, OpenCode, Kimi Code) into self-contained, embeddable HTML replays
LLM-as-a-Verifier is a general-purpose framework that provides fine-grained feedback for any agent without requiring additional training. It achieves SOTA performance across coding, robotics, and m…
AgentLab: An open-source framework for developing, testing, and benchmarking web agents on diverse tasks, designed for scalability and reproducibility.
An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.
🌎💪 BrowserGym, a Gym environment for web task automation