Lists (1)
Sort Name ascending (A-Z)
Starred repositories
(ECCV 2026): Official code for Flash-BoN: Instant Drafts for Inference-Time Scaling in Diffusion Models
Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.
Launch Claude Code with Hugging Face Inference Providers
JoyAI-Image is the unified multimodal foundation model for image understanding, text-to-image generation, and instruction-guided image editing.
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
Small efficient runtime package for quantized diffusion models
Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
FlashRT is a high-performance realtime inference engine for small-batch, latency-sensitive AI workloads. The flagship integration is production VLA control for Pi0, Pi0.5, GROOT N1.6, and Pi0-FAST.…
[IROS 2025 Best Paper Award Finalist & IEEE TRO 2026] The Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems
[ICML'26] Code and website for Self-Flow: Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis
Reviews pull requests with any OpenAI-compatible LLM
A community trust management system based on explicit vouches to participate.
Rust client for the huggingface hub aiming for minimal subset of features over `huggingface-hub` python package
TPU inference for vLLM, with unified JAX and PyTorch support.
A CLI to estimate inference memory requirements for Hugging Face models, written in Python.
Comprehensive Hugging Face Hub Library for rust
An Extensible Toolkit for Finetuning and Inference of Large Foundation Models. Large Models for All.
A family of compressed models obtained via pruning and knowledge distillation
A feature-rich command-line audio/video downloader
AI agents running research on single-GPU nanochat training automatically
Autoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton kernels.
[ICML2026] From Statics to Dynamics: Physics-Aware Image Editing with Latent Transition Priors
An implementation of 'simple diffusion: End-to-end diffusion for high resolution images' as published by Hoogeboom et al.