Stars
Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦
Lightweight SQL-based stream processing engine for IoT edge.
SkillsBench evaluates how well skills work and how effective agents are at using them.
LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
Linux Runtime Security and Forensics using eBPF
A guidance language for controlling large language models.
DSPy: The framework for programming—not prompting—language models
Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
817 structured cybersecurity skills for AI agents · Mapped to 6 frameworks: MITRE ATT&CK, NIST CSF 2.0, MITRE ATLAS, D3FEND, NIST AI RMF & MITRE F3 (Fight Fraud) · agentskills.io standard · Works w…
Spec-driven development (SDD) for AI coding assistants.
Incredibly fast JavaScript runtime, bundler, test runner, and package manager – all in one
Light, fluffy, and always free - The AWS Local Emulator alternative
The fastest, litest AI Gateway. Rust core with Python SDK. Call 100+ LLM APIs in OpenAI (or native) format with cost tracking, guardrails, load balancing, and logging [Bedrock, Azure, OpenAI, Anthr…
eBPF-based Security Observability and Runtime Enforcement
Rust-native, pattern-first stream processing engine (CEP): filters, joins, enrichment, windows—low latency on-prem & Kubernetes.
A flexible, performant and reliable search database without the AI bullshit.
Expressive statecharts and FSMs for modern Python.
VectorFlow is a high volume vector embedding pipeline that ingests raw data, transforms it into vectors and writes it to a vector DB of your choice.
A single CLAUDE.md file to improve Claude Code behavior, derived from Andrej Karpathy's observations on LLM coding pitfalls.
Breakthrough Method for Agile Ai Driven Development
The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.
Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflo…
Paella: Low-latency Model Serving with Virtualized GPU Scheduling
Simple C++ and CMake wrapper around tree-sitter.
MultiArchKernelBench: A Multi-Platform Benchmark for Kernel Generation
Production-Ready MCP Server Framework • Build, deploy & scale secure AI agent infrastructure • Includes Auth, Observability, Debugger, Telemetry & Runtime • Run real-world MCPs powering AI Agents