Stars
Official Repository of Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories
A framework for serving and evaluating LLM routers - save LLM costs without compromising quality
PinchBench is a benchmarking system for evaluating LLM models as OpenClaw coding agents. Made with 🦀 by the humans at https://kilo.ai
The repo for paper: Exploiting the Index Gradients for Optimization-Based Jailbreaking on Large Language Models.
AI agent toolkit: unified LLM API, agent loop, TUI, coding agent CLI
HyperEyes is a parallel multimodal search agent that fuses visual grounding and retrieval into a single atomic action, enabling concurrent search across multiple entities while treating inference e…
An agentic skills framework & software development methodology that works.
Heuristic Learning Blog Post
C++-based high-performance parallel environment execution engine (vectorized env) for general RL environments.
DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering
The agent that grows with you
The best-benchmarked open-source AI memory system. And it's free.
原汁原昧 Claude Code 可运行,可构建, 可调试版; 生产级工程化, 企业级可靠性; 安全无毒, 内存泄露修复
Lightweight coding agent that runs in your terminal
GlyphBanana: Advancing Precise Text Rendering Through Agentic Workflows
A Multimodal Reasoning Agent with Stateful Experiences
AI agents running research on single-GPU nanochat training automatically
OpenClaw skills for deep search — multi-source search, content extraction, and structured research reports.
Fully autonomous & self-evolving research from idea to paper. Chat an Idea. Get a Paper. 🦞
Tool-Genesis: A Task-Driven Tool Creation Benchmark for Self-Evolving Language Agent
Rewards as Labels: Revisiting RLVR from a Classification Perspective
Use Microsoft Edge's online text-to-speech service from Python WITHOUT needing Microsoft Edge or Windows or an API key