Starred repositories
Local-first code intelligence graph for MCP and CLI. Builds a persistent map of your codebase so AI coding tools read only what matters, with benchmarked context reductions on reviews and large-rep…
Visualizer for neural network, deep learning and machine learning models
Machine learning inference library for ARC EM and HS Processors
Machine learning compiler based on MLIR for Sophgo TPU.
Agent skill for beautiful, verifiable architecture, workflow, sequence, data-flow, and lifecycle diagrams—self-contained HTML with motion and crisp export.
SOTA low-bit LLM quantization (INT8/FP8/MXFP8/INT4/MXFP4/NVFP4) & sparsity; leading model compression techniques on PyTorch, TensorFlow, and ONNX Runtime
A pytorch quantization backend for optimum
🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimization tools
A single CLAUDE.md file to improve Claude Code behavior, derived from Andrej Karpathy's observations on LLM coding pitfalls.
Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.
Skills for Real Engineers. Straight from my .agents directory.
The Torch-MLIR project aims to provide first class support from the PyTorch ecosystem to the MLIR ecosystem.
A high-throughput and memory-efficient inference and serving engine for LLMs
🦸 AI 编程超能力 · 中文增强版 — superpowers(116k+ ⭐)完整汉化 + 6 个中国原创 skills,让 Claude Code / Copilot CLI / Hermes Agent / Cursor / Windsurf / Kiro / Gemini CLI 等 16 款 AI 编程工具真正会干活
A collection of pre-trained, state-of-the-art models in the ONNX format
ONNX Runtime: cross-platform, high performance ML inferencing and training accelerator
Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels
DeepEP: an efficient expert-parallel communication library
A Datacenter Scale Distributed Inference Serving Framework
LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
[MLSys 2024 Best Paper Award] AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
Development repository for the Triton language and compiler
[ICML 2023] SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
We write your reusable computer vision tools. 💜
Pre-indexed code knowledge graph, auto syncs on code changes, for Claude Code, Codex, Gemini, Cursor, OpenCode, AntiGravity, Kiro, and Hermes Agent — fewer tokens, fewer tool calls, 100% local
Bridge local AI coding agents (Claude Code, Cursor, Gemini CLI, Codex) to messaging platforms (Feishu/Lark, DingTalk, Slack, Telegram, Discord, LINE, WeChat Work). Chat with your AI dev assistant f…
Persistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with C…
omo/lazycodex: The coding agent for tokenmaxxers;the one and only agent harness for complex codebases. For your Codex, for your OpenCode