Skip to content
View wxjiao's full-sized avatar
:octocat:
Focusing
:octocat:
Focusing

Block or report wxjiao

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Official Repository of Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories

Python 54 11 Updated Aug 5, 2026

A framework for serving and evaluating LLM routers - save LLM costs without compromising quality

Python 5,322 415 Updated Aug 10, 2024

Open Frontier Intelligence

8,331 638 Updated Aug 6, 2026

npx ccusage

Rust 17,839 788 Updated Aug 10, 2026

PinchBench is a benchmarking system for evaluating LLM models as OpenClaw coding agents. Made with 🦀 by the humans at https://kilo.ai

Python 1,309 151 Updated Jul 2, 2026

Agents' Last Exam

Python 937 57 Updated Aug 5, 2026

The repo for paper: Exploiting the Index Gradients for Optimization-Based Jailbreaking on Large Language Models.

Python 15 3 Updated Dec 16, 2024

AI agent toolkit: unified LLM API, agent loop, TUI, coding agent CLI

TypeScript 86,656 10,772 Updated Aug 10, 2026

HyperEyes is a parallel multimodal search agent that fuses visual grounding and retrieval into a single atomic action, enabling concurrent search across multiple entities while treating inference e…

Python 72 Updated May 23, 2026

An agentic skills framework & software development methodology that works.

Shell 270,229 24,158 Updated Aug 8, 2026

Heuristic Learning Blog Post

Python 609 60 Updated May 25, 2026

C++-based high-performance parallel environment execution engine (vectorized env) for general RL environments.

C++ 1,494 140 Updated Jul 17, 2026

DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm

C 21,111 1,907 Updated Aug 9, 2026

MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering

Python 1,681 258 Updated Apr 24, 2026
Python 1 Updated Apr 19, 2026

The agent that grows with you

Python 228,455 44,937 Updated Aug 11, 2026

The best-benchmarked open-source AI memory system. And it's free.

Python 58,284 7,496 Updated Aug 8, 2026

原汁原昧 Claude Code 可运行,可构建, 可调试版; 生产级工程化, 企业级可靠性; 安全无毒, 内存泄露修复

TypeScript 21,956 16,480 Updated Aug 10, 2026

Lightweight coding agent that runs in your terminal

Rust 105,152 15,924 Updated Aug 11, 2026

GlyphBanana: Advancing Precise Text Rendering Through Agentic Workflows

Python 5 Updated Mar 20, 2026

A Multimodal Reasoning Agent with Stateful Experiences

Python 24 Updated Mar 31, 2026

AI agents running research on single-GPU nanochat training automatically

Python 93,592 13,289 Updated Mar 26, 2026

OpenClaw skills for deep search — multi-source search, content extraction, and structured research reports.

Python 436 34 Updated Mar 18, 2026

Fully autonomous & self-evolving research from idea to paper. Chat an Idea. Get a Paper. 🦞

Python 13,994 1,636 Updated Jul 13, 2026

Tool-Genesis: A Task-Driven Tool Creation Benchmark for Self-Evolving Language Agent

Python 7 Updated Mar 18, 2026

Rewards as Labels: Revisiting RLVR from a Classification Perspective

Python 24 Updated Jun 26, 2026

Use Microsoft Edge's online text-to-speech service from Python WITHOUT needing Microsoft Edge or Windows or an API key

Python 11,704 1,075 Updated Mar 22, 2026
Next