Highlights
- Pro
Pinned Loading
-
togethercomputer/MoA
togethercomputer/MoA PublicTogether Mixture-Of-Agents (MoA) – 65.1% on AlpacaEval with OSS models
-
-
auto-bench-audit
auto-bench-audit PublicAutomated auditing pipeline for LLM and agent benchmarks — surfaces task ambiguity, environment conflicts, and evaluation bugs.
-
alexisfox7/PRO-LONG
alexisfox7/PRO-LONG PublicProgrammatic memory for long-horizon LLM agents: the harness appends everything to one log, and the agent searches it with code. 97.4% on ARC-AGI-3 (arXiv:2607.20064)
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.