Skip to content

Popular repositories Loading

  1. CATArena CATArena Public

    CATArena is an engineering-level tournament evaluation platform for Large Language Model-driven code agents (LLM-driven code agents), based on an iterative competitive peer learning framework.

    Python 68 10

  2. PRDBench PRDBench Public

    Python 46 4

  3. amemgym amemgym Public

    Python 42 7

  4. FORTE FORTE Public

    FORTE (Full-cycle Office Real-world Task Evaluation) is a general agent benchmark for evaluating AI agents on daily office productivity across 15 corporate professions.

    Python 19 1

  5. CoreCodeBench CoreCodeBench Public

    Python 16

  6. agi-eval agi-eval Public

    Python 15

Repositories

Showing 10 of 17 repositories

People

This organization has no public members. You must be a member to see who’s a part of this organization.

Top languages

Loading…

Most used topics

Loading…