Skip to content
View woodx9's full-sized avatar
🙂
🙂

Block or report woodx9

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

NVIDIA Inference Xfer Library (NIXL)

C++ 1,184 393 Updated Aug 9, 2026

A workload for deploying LLM inference services on Kubernetes

Go 272 71 Updated Aug 8, 2026

【A common used C++ & Python DAG framework】 一个通用的、无三方依赖的、跨平台的、收录于awesome-cpp的、基于流图的并行计算框架。欢迎star & fork & 交流

C++ 2,292 388 Updated Aug 7, 2026

A Datacenter Scale Distributed Inference Serving Framework

Rust 7,713 1,421 Updated Aug 9, 2026

vLLM’s reference system for K8S-native cluster-wide deployment with community-driven performance optimization

Python 2,499 461 Updated Aug 5, 2026

A simple, open source bilingual translation extension & Greasemonkey script (一个简约、开源的 双语对照翻译扩展 & 油猴脚本)

JavaScript 11,907 559 Updated Aug 8, 2026

SGLang Omni: High-Performance Multi-Stage Pipeline Framework for Omni Models

Python 767 316 Updated Aug 9, 2026

A framework for efficient model inference with omni-modality models

Python 5,983 1,432 Updated Aug 9, 2026

TokenSpeed is a speed-of-light LLM inference engine.

Python 1,836 225 Updated Aug 9, 2026

A programmable Mixture-of-Models router for heterogeneous LLM inference

Go 5,130 805 Updated Aug 9, 2026

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python 11,078 1,679 Updated Aug 9, 2026

The official Go library for the OpenAI API

Go 3,395 344 Updated Aug 7, 2026

Milvus is a high-performance, cloud-native vector database built for scalable vector ANN search

Go 45,571 4,170 Updated Aug 9, 2026

Based on the RV32I ISA, aiming to implement the complete functions of the CPU without considering synthesis, timing, and latency.

Verilog 2 Updated Jun 20, 2025

Train speculative decoding models effortlessly and port them smoothly to SGLang serving.

Python 1,058 311 Updated Aug 9, 2026

Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM

Python 3,647 607 Updated Aug 7, 2026

FP16xINT4 LLM inference kernel that can achieve near-ideal ~4x speedups up to medium batchsizes of 16-32 tokens.

Python 1,123 90 Updated Sep 4, 2024

从零构建大模型:从预训练到RLHF的完整实践

Python 2,679 209 Updated May 20, 2026

Algorithm powering the For You feed on X

Rust 26,957 4,590 Updated May 15, 2026

a embedding infer server faster than vllm and sglang

Python 17 1 Updated Feb 10, 2026
Python 1,302 135 Updated May 20, 2026

Nano vLLM

Python 14,920 2,436 Updated Apr 26, 2026

A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.

Python 4,715 778 Updated May 17, 2026

FlashInfer: Kernel Library for LLM Serving

Python 6,133 1,249 Updated Aug 9, 2026

Fast and memory-efficient exact attention

Python 24,656 2,972 Updated Aug 9, 2026

The ultimate RAG for your monorepo. Query, understand, and edit multi-language codebases with the power of AI and knowledge graphs

Python 2,676 455 Updated Aug 9, 2026
Python 65 11 Updated Jun 19, 2024

LeetGPU Solutions

Python 128 5 Updated Oct 9, 2025

leetTriton

Python 3 Updated Sep 9, 2025

Getting Started with Triton: A Tutorial for Python Beginners

HTML 61 7 Updated Mar 26, 2026
Next