Stars
MPIE-Bench: Benchmark for multi-person character-consistent image editing under contact, with a six-axis evaluation protocol.
A 3B-active-parameter native unified multimodal model for image and video understanding, generation, and editing.
AI Agent prompts for deep paper reading — auto-routes to benchmark, methodology, or survey/opinion analysis frameworks. Works with Cursor, Claude Code, Codex, and OpenCode.
A-RAG: Agentic Retrieval-Augmented Generation via Hierarchical Retrieval Interfaces. State-of-the-art RAG framework with keyword, semantic, and chunk read tools for multi-hop QA.
[ACL 2026] WildGraphBench: Benchmarking GraphRAG with Wild-Source Corpora
Wiki Live Challenge: Challenging Deep Research Agents with Expert-Level Wikipedia Articles
DeepResearch Bench II (DRB2) is the follow-up to DeepResearch Bench, with a stronger focus on measuring the gap between deep research systems and human experts. It does so by decomposing expert-wri…
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
Self-evolving memory OS for LLM & AI Agents: ultra-persistent memory, hybrid-retrieval, and cross-task skill reuse, with 35.24% token savings