Skip to content
View AL-377's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report AL-377

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

The roadmap of long-horizon agents

955 37 Updated Aug 15, 2026

Framework for evaluating and improving agents

Python 4,262 1,571 Updated Aug 15, 2026

EdgeBench: Unveiling scaling laws of learning from real-world environments

Python 421 17 Updated Jul 17, 2026

A benchmark for evaluating AI agents on realistic business workflows

Python 199 22 Updated Aug 4, 2026

Harness for running and evaluating AI agents against RL environments

Python 235 53 Updated Aug 11, 2026

Benchmark self-evolving Agent upon realistic large-scale file workspaces

Python 55 5 Updated Aug 4, 2026

Implementation for: Towards On-Policy Data Evolution for Visual-Native Multimodal Deep Search Agents

Python 28 1 Updated May 13, 2026

Context engineering for AI Agents. Manage your LLM's context window with caching, compaction, summarization, and graceful degradation.

Python 3 1 Updated Aug 14, 2026

Trae Agent is an LLM-based agent for general purpose software engineering tasks.

Python 12,022 1,339 Updated Feb 5, 2026

🔥 A collection of the Claude Code open source

TypeScript 2,800 2,451 Updated Apr 11, 2026

The agent benchmark that scores the full stack — harness, config, and model — not just the LLM. Trace-based scoring, reliability metrics, configuration diagnostics.

Python 136 27 Updated Jul 31, 2026

The agent that grows with you

Python 230,959 45,834 Updated Aug 15, 2026

Official PyTorch implementation of "Visually-grounded Humanoid Agents"

49 1 Updated Apr 10, 2026

"百战百胜,非善之善者也;不战而屈人之兵,善之善者也。"

6 2 Updated Apr 9, 2026

毛选.skill — 让毛泽东的思维框架帮你分析问题、制定策略、看透本质。7个核心心智模型 · 10条决策启发式 · 完整表达DNA。不是复读语录,是用他的认知框架帮你看问题。

1,045 107 Updated Jul 8, 2026
Python 1,212 18 Updated Apr 27, 2026

An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.

Rust 195,048 109,118 Updated Aug 6, 2026

feishu-cli 是一个功能完整的飞书开放平台命令行工具。它将飞书文档、知识库、电子表格、消息、日历、任务等操作封装为简洁的命令行接口,核心能力是 Markdown ↔ 飞书文档双向无损转换。

Go 1,351 140 Updated Jul 31, 2026

Create beautiful slides on the web using a coding agent's frontend skills

JavaScript 27,568 2,237 Updated Jun 23, 2026

Toolathlon-Gym for testing AI agents real-world tool-use capabilities across diverse MCP servers.

Python 145 13 Updated Jul 22, 2026

Claw-Eval is an evaluation harness for evaluating LLM as agents. All tasks verified by humans.

Python 747 73 Updated Aug 15, 2026

Bash is all you need - A nano claude code–like 「agent harness」, built from 0 to 1

Python 74,293 12,031 Updated Aug 15, 2026

Create, Evaluate, and Connect AI Skills

Python 1,139 132 Updated Aug 9, 2026
Jupyter Notebook 43 2 Updated Feb 17, 2026
Python 17 1 Updated Jun 21, 2024

Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning

Python 426 49 Updated May 28, 2026
Python 25 3 Updated May 20, 2025

Yunjue Agent: A Fully Reproducible, Zero-Start In-Situ Self-Evolving Agent System for Open-Ended Tasks

Python 527 62 Updated Mar 6, 2026
HTML 37 4 Updated Mar 23, 2026
Next