A lightweight vLLM implementation built from scratch.
- 🚀 Fast offline inference - Comparable inference speeds to vLLM
- 📖 Readable codebase - Clean implementation in ~ 1,200 lines of Python code
- ⚡ Optimization Suite - Prefix caching, Tensor Parallelism, Torch compilation, CUDA graph, etc.
pip install git+https://github.com/GeeeekExplorer/nano-vllm.gitTo download the model weights manually, use the following command:
huggingface-cli download --resume-download Qwen/Qwen3-0.6B \
--local-dir ~/huggingface/Qwen3-0.6B/ \
--local-dir-use-symlinks FalseSee example.py for usage. The API mirrors vLLM's interface with minor differences in the LLM.generate method:
from nanovllm import LLM, SamplingParams
llm = LLM("/YOUR/MODEL/PATH", enforce_eager=True, tensor_parallel_size=1)
sampling_params = SamplingParams(temperature=0.6, max_tokens=256)
prompts = ["Hello, Nano-vLLM."]
outputs = llm.generate(prompts, sampling_params)
outputs[0]["text"]See bench.py for benchmark.
Test Configuration:
- Hardware: RTX 4070 Laptop (8GB)
- Model: Qwen3-0.6B
- Total Requests: 256 sequences
- Input Length: Randomly sampled between 100–1024 tokens
- Output Length: Randomly sampled between 100–1024 tokens
Performance Results:
| Inference Engine | Output Tokens | Time (s) | Throughput (tokens/s) |
|---|---|---|---|
| vLLM | 133,966 | 98.37 | 1361.84 |
| Nano-vLLM | 133,966 | 93.41 | 1434.13 |
nano-vllm/
├── pyproject.toml # 项目配置
├── example.py # 使用示例
├── bench.py # 性能基准测试
├── README.md # 项目文档
├── nanovllm/ # 核心库
│ ├── __init__.py # 导出 LLM、SamplingParams
│ ├── config.py # 配置类
│ ├── llm.py # LLM 主类(继承 LLMEngine)
│ ├── sampling_params.py # 采样参数
│ ├── engine/ # 推理引擎核心
│ │ ├── llm_engine.py # 引擎主类
│ │ ├── model_runner.py # 模型执行器(处理CUDA/分布式)
│ │ ├── scheduler.py # 请求调度器
│ │ ├── sequence.py # 序列数据结构
│ │ └── block_manager.py # KV缓存块管理
│ ├── models/ # 模型实现
│ │ └── qwen3.py # Qwen3 模型实现
│ ├── layers/ # 神经网络层
│ │ ├── attention.py # 注意力层
│ │ ├── linear.py # 线性层(支持并行)
│ │ ├── activation.py # 激活函数
│ │ ├── layernorm.py # 层归一化
│ │ ├── rotary_embedding.py # 旋转位置编码
│ │ ├── embed_head.py # 嵌入和输出头
│ │ └── sampler.py # 采样器
│ └── utils/ # 工具函数
│ ├── context.py # 执行上下文管理
│ └── loader.py # 模型权重加载
└── assets/ # 资源文件时间轴:
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
初始化:
[Rank 0] ModelRunner ✓ (CUDA Graph录制)
[Rank 1+] ModelRunner ✓ (若tensor_parallel_size > 1)
Step 1(Prefill - Req1):
Scheduler: 从 waiting 取 Req1
Model: 处理 [token_ids_req1] (长度: 100)
Sampler: 生成第101个token
Update: Req1 进入 running
Step 2(Prefill - Req2):
Scheduler: 从 waiting 取 Req2
Model: 处理 [token_ids_req2] (长度: 50)
Sampler: 生成第51个token
Update: Req2 进入 running
Step 3(Decode - Req1, Req2):
Scheduler: 从 running 取 [Req1, Req2]
Model: 处理 [last_token_req1, last_token_req2] (每个batch 1个)
Sampler: 采样 2 个 token
Update: Req1已生成56/256, Req2已生成2/256
Step 4(Decode - 继续):
......循环直到所有请求达到max_tokens或EOS
完成:
返回: [{"text": "...", "token_ids": [...]}, {"text": "...", ...}]