Popular repositories Loading
Repositories
Showing 10 of 17 repositories
- FORTE Public
FORTE (Full-cycle Office Real-world Task Evaluation) is a general agent benchmark for evaluating AI agents on daily office productivity across 15 corporate professions.
- DailyReport Public
DailyReport: An Open-ended Benchmark for Evaluating Search Agents on Daily Search Tasks
- PRDBench Public
- Repo-of-AgentEscapeBench Public
AgentEscapeBench: A benchmark for evaluating out-of-domain tool-grounded reasoning in LLM agents via procedurally-generated escape room challenges with DAG-structured dependencies.
- SIGHT Public
- amemgym Public
- BrowseComp-ZH-revised Public
- tau2-bench-revised Public
People
This organization has no public members. You must be a member to see who’s a part of this organization.
Top languages
Loading…
Most used topics
Loading…