-
The Hong Kong University of Science and Technology
- Hong Kong SAR, China
Stars
Lightweight, open-source AI agent for your tools, chats, and workflows.
Synchronize Codex session provider metadata across rollout files and SQLite state.
Open source Ghostty-based macOS terminal with vertical tabs and notifications for AI coding agents. Built for multitasking, organization, and programmability.
Wrap Antigravity, ChatGPT Codex, Claude Code, Grok Build as an OpenAI/Gemini/Claude/Codex compatible API service, allowing you to enjoy the free Gemini 3.1 Pro, GPT 5.5, Grok 4.3, Claude model thro…
Use Garry Tan's exact Claude Code setup: 23 opinionated tools that serve as CEO, Designer, Eng Manager, Release Manager, Doc Engineer, and QA
ARIS ⚔️ (Auto-Research-In-Sleep) — Lightweight Markdown-only skills for autonomous ML research: cross-model review loops, idea discovery, and experiment automation. No framework, no lock-in — works…
LLM 驱动的多市场股票智能分析系统:多源行情、实时新闻、决策看板与自动推送,支持零成本定时运行。 LLM-powered multi-market stock analysis system with multi-source market data, real-time news, decision dashboard, automated notifications, and cost…
A live reading list for LLM data synthesis (Updated to July, 2025).
Segment Policy Optimization: Effective Segment-Level Credit Assignment in RL for Large Language Models
The agent that grows with you
A benchmark for LLMs on complicated tasks in the terminal
PinchBench is a benchmarking system for evaluating LLM models as OpenClaw coding agents. Made with 🦀 by the humans at https://kilo.ai
Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
SkillsBench evaluates how well skills work and how effective agents are at using them.
Framework for evaluating and improving agents
Claw-Eval is an evaluation harness for evaluating LLM as agents. All tasks verified by humans.
One local control plane for every AI agent: route across models, fuse new capabilities, orchestrate tools, and stay fully in control.
MCPMark is a comprehensive, stress-testing MCP benchmark designed to evaluate model and agent capabilities in real-world MCP use.
Salesforce Enterprise Deep Research
MCP-Universe is a comprehensive framework designed for RL training, benchmarking, and developing AI agents for general tool-use.
A Survey of Reinforcement Learning for Large Reasoning Models
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
基于多智能体LLM的中文金融交易框架 - TradingAgents中文增强版
Tongyi Deep Research, the Leading Open-source Deep Research Agent