Starred repositories
Orca is the ADE for working with a fleet of parallel agents. Run any coding agent with your own subscription. Available on desktop, mobile and remote runtime.
Automatic C-Go Bindings Generator for Go Programming Language
Secure, fast, efficient, battle-tested at Alibaba's scale. Hybrid architecture code review tool: deterministic pipelines + LLM Agent, precise line-level comments, built-in multi-language ruleset (N…
GitNexus: The Zero-Server Code Intelligence Engine
Claude Code skills for Kubernetes platform engineering, GitOps, and Helm chart management
Kubernetes Skill for Claude Code and Codex. LLMs hallucinate a lot with K8s - KubeShark fixes this. It eliminates hallucinations and grounds your Kubernetes, Helm etc official best practices.
Production-grade engineering skills for AI coding agents.
The Python implementation of Connect: Protobuf RPC that works.
Fast CUDA matrix multiplication from scratch
A C++ header-only HTTP/HTTPS server and client library
Transformer Explained Visually: Learn How LLM Transformer Models Work with Interactive Visualization
Diffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference in pure C/C++
SGLang is a high-performance serving framework for large language models and multimodal models.
Port of OpenAI's Whisper model in C/C++
GGUF Quantization support for native ComfyUI models
A really basic thread-safe progress bar for Golang applications
Performance-portable, length-agnostic SIMD with runtime dispatch
Machine Learning Containers for NVIDIA Jetson and JetPack-L4T
The TinyLlama project is an open endeavor to pretrain a 1.1B Llama model on 3 trillion tokens.
FlatBuffers: Memory Efficient Serialization Library