Skip to content
View JackyCSer's full-sized avatar
  • Hangzhou, China

Block or report JackyCSer

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

ASI-Bench

Python 10 Updated Aug 10, 2026

Orca is the ADE for working with a fleet of parallel agents. Run any coding agent with your own subscription. Available on desktop, mobile and VPS.

TypeScript 41,608 2,910 Updated Aug 10, 2026

Lightweight loop engineering state kernel for long-running AI agent teams. Agent-loop agnostic across Codex, Claude Code, and other coding agents, with durable goals, quota-aware auto-wake, executa…

Python 3,910 314 Updated Aug 10, 2026
Python 45 4 Updated Jun 29, 2026

Measuring and evolving with the frontier of agent work

Python 464 370 Updated Aug 10, 2026

🤯 LobeHub is your Chief Agent Operator, organizing your agents into 7×24 operations by hiring, scheduling, and reporting on your entire AI team.

TypeScript 81,469 15,788 Updated Aug 10, 2026
Python 263 30 Updated Jul 16, 2026

Samaya AI's FrontierFinance Benchmark Grader

Python 18 1 Updated Jul 16, 2026

ReactBench is an evaluation for coding agents on realistic React work

JavaScript 241 9 Updated Jul 27, 2026

A Kotlin software-engineering benchmark for evaluating coding agents.

Python 49 6 Updated Aug 6, 2026

Code for paper AIDABench: AI Data Analytics Benchmark.

Python 109 3 Updated Jun 24, 2026

WideSearch: Benchmarking Agentic Broad Info-Seeking

Python 150 18 Updated Oct 9, 2025

Terminal-Bench Science: Evaluating AI Agents on Complex Real-World Scientific Workflows in the Terminal

Python 241 211 Updated Aug 10, 2026

Real-world buy-side tasks for evaluating agents on investment research. China-centered, globally aware (A-share · HK · US · macro). Public set.

10 Updated Jul 1, 2026

A benchmark for evaluating AI agents on realistic business workflows

Python 185 20 Updated Aug 4, 2026

The fastest way to put Volcengine Ark in your terminal and your AI agent — go from prompt to generated media, multimodal answer, or deployed endpoint in a single command, no API glue code.

Python 93 7 Updated Aug 6, 2026

Evals Harness for $OneMillion-Bench

Python 47 5 Updated Jul 30, 2026

Framework for evaluating and improving agents

Python 4,077 1,510 Updated Aug 10, 2026

SWE-Marathon: an ultra long-horizon SWE benchmark

Rust 130 26 Updated Aug 10, 2026

Agents' Last Exam

Python 935 57 Updated Aug 5, 2026

Harbor is a framework for running agent evaluations and creating and using RL environments.

Python 1 Updated Jul 10, 2026

Harness for running and evaluating AI agents against RL environments

Python 232 53 Updated Aug 10, 2026

Benchmark Everything Everywhere All at Once, a fully autonomous agentic system for benchmark construction and customization.

Python 28 2 Updated Jun 10, 2026

Anonymous Github is a proxy server to support anonymous browsing of Github repositories for open-science code and data.

TypeScript 2,165 86 Updated Aug 6, 2026

Goal-driven harness experiments on real agent runtimes — a local, stdlib-only CLI workbench with Copilot + Auto modes and evidence-first reporting.

Python 2 Updated Jul 15, 2026

Pier is a Harbor fork built for DeepSWE, with stronger support for CLI agents in air-gapped (no-internet) tasks and more faithful, consistent agent trajectories

Python 141 40 Updated Aug 6, 2026

Measuring frontier coding agents on original, long-horizon engineering tasks

Python 1,341 87 Updated Aug 6, 2026

A coding agent framework, that works on its own codebase.

Python 380 62 Updated Apr 23, 2025

Benchmarking LLMs on Real-World Software Verification in Lean 4

Python 11 3 Updated Aug 10, 2026

Lightweight coding agent that runs in your terminal

Rust 105,137 15,917 Updated Aug 10, 2026
Next