-
18:05
(UTC -12:00)
Lists (19)
Sort Name ascending (A-Z)
Agent
agent benchmark
agent memory
Agentic RL
coding agent
dLLM
EAI
Generation
LLM
LLM Reasoning
low vision
low vision listmamba
mamba listmedical image
MLLM
Mutil Model Agent
paper list
paper writing
RLHF
World Model
Stars
Infinite Worlds with Versatile Interactions
VibeGame: Vibe Your Dream Game -- An open-source self-evolving multi-agent framework with an AI-Native game engine that turns your natural language into a fully playable 2D web game and edit it any…
A curated list of plugins, skills, MCP servers, patch/profile layers, orchestrators & UIs for DeepSeek Harness (DSH). Visualization · PPT · Coding · Agents · Loops (auto-research) and more. #dsh
Official implementation of EVOKE: Endless Interactive World with Bounded State and Long-Horizon Supervision. A three-step, CFG-free interactive world model. SOTA on WBench.
A curated, continuously updated reading list, paper blogs, and resources for World Action Models (WAMs) in embodied AI.
The open-source design agent and harness, better than Claude Design on academic communication artifacts production. This DesignHarness can also be used with any coding harness you like ( Codex/Clau…
A benchmark for evaluating AI agents on realistic business workflows
DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation
💻 SETA: Scaling Environments for Terminal Agents - Environments
Toonflow 是开源一站式 AI 短剧创作工具,将小说、剧本快速转化为动画短剧。集成 AI 编剧、智能分镜、角色与视频生成,跨平台桌面端轻量部署,助力创作者低成本批量产出视觉内容。Toonflow is an open-source AI tool that turns stories and scripts into animated short dramas. Features AI…
The lightweight framework for building agents
Benchmarking Agents on Real-World Terminal Tasks
Code for the paper "Self-Compacting Language Model Agents"
The roadmap of long-horizon agents
Infinite Interactive World Rollout on a Single Desktop GPU
🔥 Quo Vadis, World Modeling? Towards Interactive World Proxies for Continually Improving Agents
Qwen-CUA: Native Computer Use for (Almost) Everything — a screenshot-driven agent that operates computers with keyboard and mouse, jointly developed by the Qwen Team and XLang Lab.
🐧 Harness for RSI. Let AI Build AI. Everything is Transparent.
The long-horizon computer-use harness. Run AI agents across desktop apps and the CLI for extended periods while preserving task state and making reliable progress on complex workflows. Features fre…
Long Horizon Terminal Benchmark with Dense Reward Grading
A benchmark for general-purpose terminal-use agents.
WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces