Stars
Awesome AI Memory | LLM Memory | A curated knowledge base on AI memory for LLMs and agents, covering long-term memory, reasoning, retrieval, and memory-native system design. Awesome-AI-Memory 是一个 集…
DeepSeek Harness: Everything is a Plugin.
Omnigent is an open-source AI agent framework and meta-harness: orchestrate Claude Code, Codex, Cursor, Pi, and custom agents — swap harnesses without rewriting, enforce policies and sandboxing, an…
TencentDB Agent Memory is a team-level memory hub for AI Agents — turning conversations, docs, and code into four reusable memory assets (Chat Memory, Skill, LLM-Wiki, Code-Graph) that are governed…
Local-first AI agent workspace for coding, writing, design, research, and automation — one runtime for desktop GUI and TUI.
A framework for running evals against small (and large) models
Collab with OpenAI. A benchmark and harness for finding and exploiting smart contract bugs
Privacy-first, local-only personal OS — calendar, notes, kanban, spreadsheets, AI terminal, and 30+ widgets. All data stays on your device. React 19 + TypeScript + IndexedDB.
An evaluation and evolution tool for Agent Skills.
SkillOpt is a text-space optimizer that trains reusable natural-language skills for frozen LLM agents through trajectory-driven edits, validation-gated updates, and deployable best_skill.md artifacts.
AI turns documents or topics into real, native PowerPoint decks—with native shapes, transitions and animations, data-backed charts and tables on demand, audio narration from speaker notes, and supp…
🎨 Best DeepSeek Harness Design Plugin. The open-source Claude Design alternative. 🖥️ Local-first desktop app. 🖼️ Your coding agent becomes the design engine: prototypes, landing pages, dashboards, …
AST-based outline director for human-centered AI presentation workflows.
Unified benchmark for evaluating conversational memory and RAG across multiple datasets
Self-organizing AI second brain for Obsidian + Claude Code. Drop any source and Claude reads, links, and files it into one connected knowledge graph of plain Markdown you own. AI note-taking, perso…
SkillsBench evaluates how well skills work and how effective agents are at using them.
DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms
EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions
Open source code for ICLR 2026 Paper: Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions
One portable memory layer for every AI agent: local-first, Markdown-native, user-owned, and self-evolving across apps, tools, and workflows.
Hindsight: Agent Memory That Learns
Build Real-Time Knowledge Graphs for AI Agents
[NeurIPS 2025] Open-source Multi-agent Poster Generation from Papers
HY-SOAR:Self-Correction for Optimal Alignment and Refinement in Diffusion Models
[CVPR'26 Highlight] AMB3R: Accurate Feed-forward Metric-scale 3D Reconstruction with Backend
Performance analysis of predictive (alpha) stock factors