-
Lehigh University
- Bethlehem, US
- https://bbsngg.github.io/
Lists (1)
Sort Name ascending (A-Z)
Starred repositories
HERO = Hashing Β· Edge cases Β· Rubrics Β· Overbuild β the four shapes coding agents over-defend in. A paste-in contract that stops them. Works with Claude Code, Codex, Antigravity, Cursor, Copilot, Wβ¦
Open Science Desktop β local-first, model-agnostic AI research workbench for macOS, Windows & Linux. Open-source Claude Science desktop alternative built on Tauri + MCP + agent skills.
A debugging framework for agentic AI systems: diagnose failures, attribute root causes, recover with evidence, and validate fixes through reruns.
π¦ Just talk to your agent β it learns and EVOLVES π§¬.
A general framework for distilling human-created multimodal resources into reusable, executable skills that AI agents can browse, compose, and run, validated across diverse domains including web, Pβ¦
Reference code for the Meta-Harness paper.
The roadmap of long-horizon agents
This is the repo for the paper TerminalTraj: Large-Scale Terminal Agentic Trajectory Generation from Dockerized Environments
The open-source AI workbench for scientific research
Don't trust an autoresearch paper at face value. Reviewer-side integrity forensics (self-consistency + fabrication), deterministic verdict. 61 signals: 46 integrity hack-patterns (families AβH, verβ¦
A benchmark for evaluating AI agents on frontier ultra long-horizon auto research tasks.
Repository for TRACE: Capability-Targeted Agentic Training
[Survey] A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems
Code for the multi-agent computer use project.
Open-World Self-Evolution for LLM Agents β agents that build both their skills and their own verification signals from scratch, with no target-task supervision. (Code coming soon.)
One dashboard. An entire research team.
π₯ A Survey on AI Auto-Research
π€ ml-intern: an open-source ML engineer that reads papers, trains models, and ships ML models
π OpenHands: AI-Driven Development
AIDE: an LLM agent for machine learning engineering - the research Weco grew out of. Referenced in OpenAI MLE-bench.
Dependency-Aware Structural Retrieval for Massive Agent Skills
AIRS-Bench: an AI Research Science benchmark for quantifying the end-to-end AI research abilities of LLM agents
Open-source autoresearch powered by autonomous coding agents. Run Claude Code, OpenCode, and Codex with grading, shared knowledge, and multi-agent evolution. Accepted at COLM 2026.
Now, Stronger AI Pushes Frontiers, Stronger Our Shared Future.
Dr. Claw plugin for Claude Code β AI research pipeline