Highlights
- Pro
Stars
DFlash: Block Diffusion for Flash Speculative Decoding
🟣 LLMs interview questions and answers to help you prepare for your next machine learning and data science interview in 2026.
LLM algorithm practice lab with theory, solutions, and test cases.《大模型算法与系统教程》面向大模型入门到进阶的算法实战教程,覆盖原理讲解、答案解析、测试用例与 CUDA/Triton 实战。
LLM notes, including model inference, transformer model structure, and llm framework code analysis notes.
📑 PageIndex: Document Index for Vectorless, Reasoning-based RAG
OpenTelemetry Instrumentation for AI Observability
Scan your Rust crate for semver violations.
100+ LLM interview questions with answers.
NVIDIA Isaac GR00T N1.7 - A Foundation Model for Generalist Robots.
openvla / openvla
Forked from TRI-ML/prismatic-vlmsOpenVLA: An open-source vision-language-action model for robotic manipulation.
Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.
slime is an LLM post-training framework for RL Scaling.
τ-Bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
High-performance Rust benchmark client for vLLM serving endpoints.
🧠 Train a 64M-parameter LLM from scratch in just 2h!
A programmable Mixture-of-Models router for heterogeneous LLM inference
KAI Scheduler is an open source Kubernetes Native scheduler for AI workloads at large scale
《动手学大模型Dive into LLMs》系列编程实践教程
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
vLLM’s reference system for K8S-native cluster-wide deployment with community-driven performance optimization
A workload for deploying LLM inference services on Kubernetes
LeaderWorkerSet: An API for deploying a group of pods as a unit of replication
Graphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini C…
Symphony turns project work into isolated, autonomous implementation runs, allowing teams to manage work instead of supervising coding agents.
Proxy based Redis cluster solution supporting pipeline and scaling dynamically