Skip to content
View bbsngg's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report bbsngg

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Starred repositories

Showing results

HERO = Hashing Β· Edge cases Β· Rubrics Β· Overbuild β€” the four shapes coding agents over-defend in. A paste-in contract that stops them. Works with Claude Code, Codex, Antigravity, Cursor, Copilot, W…

Markdown 199 6 Updated Aug 18, 2026

Open Science Desktop β€” local-first, model-agnostic AI research workbench for macOS, Windows & Linux. Open-source Claude Science desktop alternative built on Tauri + MCP + agent skills.

TypeScript 1,391 152 Updated Aug 18, 2026

A debugging framework for agentic AI systems: diagnose failures, attribute root causes, recover with evidence, and validate fixes through reruns.

Python 44 5 Updated Aug 18, 2026

🦞 Just talk to your agent β€” it learns and EVOLVES 🧬.

Python 3,485 454 Updated Jun 7, 2026

A general framework for distilling human-created multimodal resources into reusable, executable skills that AI agents can browse, compose, and run, validated across diverse domains including web, P…

Python 478 56 Updated Jul 17, 2026

Reference code for the Meta-Harness paper.

Python 1,426 138 Updated Jul 11, 2026

The roadmap of long-horizon agents

967 38 Updated Aug 15, 2026

Training terminal-agents

Python 281 40 Updated Aug 13, 2026

This is the repo for the paper TerminalTraj: Large-Scale Terminal Agentic Trajectory Generation from Dockerized Environments

144 1 Updated Jul 21, 2026

The open-source AI workbench for scientific research

TypeScript 3,261 445 Updated Aug 18, 2026

Don't trust an autoresearch paper at face value. Reviewer-side integrity forensics (self-consistency + fabrication), deterministic verdict. 61 signals: 46 integrity hack-patterns (families A–H, ver…

Python 138 8 Updated Aug 18, 2026

A benchmark for evaluating AI agents on frontier ultra long-horizon auto research tasks.

Python 161 18 Updated Jun 17, 2026

Repository for TRACE: Capability-Targeted Agentic Training

Python 115 12 Updated Jul 12, 2026

[Survey] A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems

2,445 186 Updated May 16, 2026
Python 101 12 Updated Mar 30, 2026

Code for the multi-agent computer use project.

Python 21 3 Updated Jul 3, 2026

Open-World Self-Evolution for LLM Agents β€” agents that build both their skills and their own verification signals from scratch, with no target-task supervision. (Code coming soon.)

86 4 Updated Jun 8, 2026

One dashboard. An entire research team.

Python 1,354 67 Updated Jun 15, 2026

πŸ”₯ A Survey on AI Auto-Research

HTML 490 34 Updated Jul 27, 2026

πŸ€— ml-intern: an open-source ML engineer that reads papers, trains models, and ships ML models

Python 10,738 1,172 Updated Jul 30, 2026

πŸ™Œ OpenHands: AI-Driven Development

TypeScript 84,422 10,986 Updated Aug 18, 2026

AIDE: an LLM agent for machine learning engineering - the research Weco grew out of. Referenced in OpenAI MLE-bench.

Python 1,480 221 Updated Aug 17, 2026
Python 655 48 Updated Jul 28, 2026

Dependency-Aware Structural Retrieval for Massive Agent Skills

Python 201 25 Updated Aug 17, 2026

AIRS-Bench: an AI Research Science benchmark for quantifying the end-to-end AI research abilities of LLM agents

Python 110 9 Updated May 5, 2026

Open-source autoresearch powered by autonomous coding agents. Run Claude Code, OpenCode, and Codex with grading, shared knowledge, and multi-agent evolution. Accepted at COLM 2026.

Python 897 118 Updated Aug 15, 2026

Now, Stronger AI Pushes Frontiers, Stronger Our Shared Future.

TypeScript 3,277 330 Updated Jun 28, 2026

Dr. Claw plugin for Claude Code β€” AI research pipeline

TeX 5 Updated May 20, 2026
Next