Lists (13)
Sort Name ascending (A-Z)
Starred repositories
An AI-native office suite for macOS and Windows: word processor, spreadsheet, presentations, and PDF.
AgentENV (AENV) is a distributed platform for running agent environments at scale.
JuiceFS is a distributed POSIX file system built on top of Redis and S3.
Faster attention kernels for serving TML's Inkling model on vLLM. 2.7x over the shipping path on H100, and the only implementation that runs on A100.
OpenMinis — The AI Agent app across platforms. Fully free and open source.
Argilla is a collaboration tool for AI engineers and domain experts to build high-quality datasets
Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.
Fast, efficient, battle-tested at Alibaba's scale. Hybrid architecture code review tool: deterministic pipelines + LLM Agent, precise line-level comments, built-in multi-language ruleset (NPE, thre…
SWE-Marathon: an ultra long-horizon SWE benchmark
Removes 20+ patterns of AI slop from any piece of writing.
Rebuild the object in a reference image as a code-only, procedural, quality-gated, animation-ready Three.js model. Token-efficient image-to-3D.
Write HTML. Render video. Built for agents.
A library of creative canvas components. Real HTML with WebGL effects running over it. React, Vue, Svelte, vanilla.
Your agent writes bad React. This catches it
SpaceXAI's coding agent harness and TUI. Fullscreen, mouse interactive, extensible.
RushDB is a graph + vector database and memory layer for AI agents. Push any JSON, get typed, searchable, relationship-aware records back — no schema, no migrations. Built on Neo4j.
Claude Code plugin for creating LLM-as-a-judge evaluators using the Plurai platform
Simple CLI tool for deterministic routing of queries between local and hosted LLM models
Let Pi control your apps on MacOS & Windows
A collection of agent skills that help with various parts of building a great interface. From animation and UI polish to accessibility and product writing.
Long Horizon Terminal Benchmark with Dense Reward Grading