Lists (1)
Sort Name ascending (A-Z)
Stars
A reference repository for building Agent Systems with the EXAONE foundation model. It provides code examples, design patterns, best practices, real-world use cases, and A-to-Z learning materials f…
Tool that just makes your open source project better using LLM agents
The lightweight framework for building agents
An in-the-wild benchmark for AI agents in the OpenClaw Environment.
🌍 AppWorld: A Controllable World of Apps and People for Benchmarking Function Calling and Interactive Coding Agent, ACL'24 Best Resource Paper.
[NeurIPS 2022] 🛒WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
Framework for evaluating and improving agents
SkillsBench evaluates how well skills work and how effective agents are at using them.
Real-time global intelligence dashboard. AI-powered news aggregation, geopolitical monitoring, and infrastructure tracking in a unified situational awareness interface
Vero: An Open RL Recipe for General Visual Reasoning
Claw-Eval is an evaluation harness for evaluating LLM as agents. All tasks verified by humans.
Browser automation CLI for AI agents
Mount Hugging Face Buckets and repos as local filesystems. No download, no copy, no waiting.
AI agent toolkit: unified LLM API, agent loop, TUI, coding agent CLI
[NeurIPS 2025] The official implementation of "KL Penalty Control via Perturbation for Direct Preference Optimization"
Zero Bubble Pipeline Parallelism
OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis
chrome & firefox extension to chat with webpages: local llms
Convert Word documents to beautiful Markdown. Via command line or in your browser.
Training library for Megatron-based models with bidirectional Hugging Face conversion capability
An open-source RAG-based tool for chatting with your documents.
[NeurIPS 2025 Spotlight] Reasoning Environments for Reinforcement Learning with Verifiable Rewards
Atropos is a Language Model Reinforcement Learning Environments framework for collecting and evaluating LLM trajectories through diverse environments