Starred repositories
AI agent toolkit: unified LLM API, agent loop, TUI, coding agent CLI
A runtime substrate that turns an agent's execution into a reversible, Git-like trace, so meta-agents can observe, fork, replay, and revert any run. Couples agent and environments in a copy-on-writ…
A self-improving RLM agent for coding workflows and long-running autonomous tasks.
Agent workspace built on Cloudflare Workers for creating documents, building apps, and running agents with your company’s context and systems.
⌥ AI Coding agent for the terminal — hash-anchored edits, optimized tool harness, LSP, Python, browser, subagents, and more
An agentic operating system where the kernel is controlled directly by Claude
A real-time procedural snow rendering demo built with WebGPU, Babylon.js and hand-written WGSL. Features GPU-generated terrain, snow deformation, procedural characters, cloth, surf wakes, water spe…
Open Science is an open-source, local-first, model-agnostic AI research workbench for scientific discovery.
AgentENV (AENV) is a distributed platform for running agent environments at scale.
A GEPA proposer with less overfitting and helpful parameters.
TokenSpeed is a speed-of-light LLM inference engine.
General plug-and-play inference library for Recursive Language Models (RLMs), supporting various sandboxes.
[ICML 2026] Materials Discovery Environment: Benchmarking Agentic systems for Closed-Loop Materials Discovery
LLM-as-a-Verifier is a general-purpose framework that provides fine-grained feedback for any agent without requiring additional training. It achieves SOTA performance across coding, robotics, and m…
Components crafted for Design Engineers. Styled using Tailwind CSS, fully compatible with Shadcn, and easy to integrate—just copy and paste. MIT 🤌
Automatic Agent for DFT calculation setup and benchmarking
NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers?
BEDC: Binary Emission Discovery Calculus (mathlib-free Lean 4 + LaTeX paper)
Turn any document or a whole zip into an interactive knowledge graph, using a self-hosted Qwen3.6-35B-A3B-MTP on a single NVIDIA L4
A benchmark for evaluating AI agents on frontier ultra long-horizon auto research tasks.
DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
The 100 line AI agent that solves GitHub issues or helps you in your command line. Radically simple, no huge configs, no giant monorepo—but scores >74% on SWE-bench verified!
Fast and memory-efficient classical machine learning operators
Winner 🏆 (Agent-only) MLSys 2026 - FlashInfer AI Kernel Generation Contest for the DeepSeek Sparse Attention (DSA) track with an average speedup of 34.93x
A high-performance toolkit for atomistic simulations in JAX.