Skip to content
View JackyCSer's full-sized avatar
  • Hangzhou, China

Block or report JackyCSer

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

DeepSeek Harness: Everything is a Plugin.

TypeScript 87,052 7,735 Updated Aug 13, 2026
Python 4 Updated Aug 11, 2026

9.9 元豆包API复刻 Claude Science

Python 316 37 Updated Aug 14, 2026

Orca is the ADE for working with a fleet of parallel agents. Run any coding agent with your own subscription. Available on desktop, mobile and VPS.

TypeScript 45,342 3,165 Updated Aug 14, 2026

Lightweight loop engineering state kernel for long-running AI agent teams. Agent-loop agnostic across Codex, Claude Code, and other coding agents, with durable goals, quota-aware auto-wake, executa…

Python 4,663 403 Updated Aug 14, 2026
Python 46 4 Updated Jun 29, 2026

Measuring and evolving with the frontier of agent work

Python 496 374 Updated Aug 13, 2026

🤯 LobeHub is your Chief Agent Operator, organizing your agents into 7×24 operations by hiring, scheduling, and reporting on your entire AI team.

TypeScript 81,682 15,809 Updated Aug 14, 2026

Community plugin to control Blender 3D with any LLM of your choice

Python 25,839 2,463 Updated Aug 11, 2026
Python 266 30 Updated Jul 16, 2026

Samaya AI's FrontierFinance Benchmark Grader

Python 18 1 Updated Jul 16, 2026

ReactBench is an evaluation for coding agents on realistic React work

JavaScript 244 10 Updated Jul 27, 2026

A Kotlin software-engineering benchmark for evaluating coding agents.

Python 49 6 Updated Aug 13, 2026

Code for paper AIDABench: AI Data Analytics Benchmark.

Python 110 3 Updated Jun 24, 2026

WideSearch: Benchmarking Agentic Broad Info-Seeking

Python 151 18 Updated Oct 9, 2025

Terminal-Bench-Science: Evaluating AI Agents on Complex Real-World Scientific Workflows in the Terminal

Python 246 233 Updated Aug 14, 2026

Real-world buy-side tasks for evaluating agents on investment research. China-centered, globally aware (A-share · HK · US · macro). Public set.

11 Updated Jul 1, 2026

A benchmark for evaluating AI agents on realistic business workflows

Python 196 21 Updated Aug 4, 2026

The fastest way to put Volcengine Ark in your terminal and your AI agent — go from prompt to generated media, multimodal answer, or deployed endpoint in a single command, no API glue code.

Python 95 7 Updated Aug 14, 2026

Evals Harness for $OneMillion-Bench

Python 47 5 Updated Jul 30, 2026

Framework for evaluating and improving agents

Python 4,237 1,561 Updated Aug 13, 2026

SWE-Marathon: an ultra long-horizon SWE benchmark

Rust 135 26 Updated Aug 14, 2026

Agents' Last Exam

Python 949 60 Updated Aug 14, 2026

Harbor is a framework for running agent evaluations and creating and using RL environments.

Python 1 1 Updated Aug 13, 2026

Harness for running and evaluating AI agents against RL environments

Python 233 53 Updated Aug 11, 2026

Benchmark Everything Everywhere All at Once, a fully autonomous agentic system for benchmark construction and customization.

Python 28 2 Updated Jun 10, 2026

Anonymous Github is a proxy server to support anonymous browsing of Github repositories for open-science code and data.

TypeScript 2,171 89 Updated Aug 10, 2026

Goal-driven harness experiments on real agent runtimes — a local, stdlib-only CLI workbench with Copilot + Auto modes and evidence-first reporting.

Python 2 Updated Jul 15, 2026

Pier is a Harbor fork built for DeepSWE, with stronger support for CLI agents in air-gapped (no-internet) tasks and more faithful, consistent agent trajectories

Python 146 43 Updated Aug 13, 2026
Next