Skip to content

About

Nano vLLM

Resources

Stars

0 stars

Watchers

0 watching

Forks

 
 

Latest commit

 

History

74 Commits

Folders and files

Repository files navigation

GeeeekExplorer%2Fnano-vllm | Trendshift

Nano-vLLM

A lightweight vLLM implementation built from scratch.

Key Features

  • 🚀 Fast offline inference - Comparable inference speeds to vLLM
  • 📖 Readable codebase - Clean implementation in ~ 1,200 lines of Python code
  • ⚡ Optimization Suite - Prefix caching, Tensor Parallelism, Torch compilation, CUDA graph, etc.

Installation

pip install git+https://github.com/GeeeekExplorer/nano-vllm.git

Model Download

To download the model weights manually, use the following command:

huggingface-cli download --resume-download Qwen/Qwen3-0.6B \
  --local-dir ~/huggingface/Qwen3-0.6B/ \
  --local-dir-use-symlinks False

Quick Start

See example.py for usage. The API mirrors vLLM's interface with minor differences in the LLM.generate method:

from nanovllm import LLM, SamplingParams
llm = LLM("/YOUR/MODEL/PATH", enforce_eager=True, tensor_parallel_size=1)
sampling_params = SamplingParams(temperature=0.6, max_tokens=256)
prompts = ["Hello, Nano-vLLM."]
outputs = llm.generate(prompts, sampling_params)
outputs[0]["text"]

Benchmark

See bench.py for benchmark.

Test Configuration:

  • Hardware: RTX 4070 Laptop (8GB)
  • Model: Qwen3-0.6B
  • Total Requests: 256 sequences
  • Input Length: Randomly sampled between 100–1024 tokens
  • Output Length: Randomly sampled between 100–1024 tokens

Performance Results:

Inference Engine Output Tokens Time (s) Throughput (tokens/s)
vLLM 133,966 98.37 1361.84
Nano-vLLM 133,966 93.41 1434.13

Code directory description

nano-vllm/
├── pyproject.toml              # 项目配置
├── example.py                  # 使用示例
├── bench.py                    # 性能基准测试
├── README.md                   # 项目文档
├── nanovllm/                   # 核心库
│   ├── __init__.py             # 导出 LLM、SamplingParams
│   ├── config.py               # 配置类
│   ├── llm.py                  # LLM 主类(继承 LLMEngine)
│   ├── sampling_params.py      # 采样参数
│   ├── engine/                 # 推理引擎核心
│   │   ├── llm_engine.py       # 引擎主类
│   │   ├── model_runner.py     # 模型执行器(处理CUDA/分布式)
│   │   ├── scheduler.py        # 请求调度器
│   │   ├── sequence.py         # 序列数据结构
│   │   └── block_manager.py    # KV缓存块管理
│   ├── models/                 # 模型实现
│   │   └── qwen3.py            # Qwen3 模型实现
│   ├── layers/                 # 神经网络层
│   │   ├── attention.py        # 注意力层
│   │   ├── linear.py           # 线性层(支持并行)
│   │   ├── activation.py       # 激活函数
│   │   ├── layernorm.py        # 层归一化
│   │   ├── rotary_embedding.py # 旋转位置编码
│   │   ├── embed_head.py       # 嵌入和输出头
│   │   └── sampler.py          # 采样器
│   └── utils/                  # 工具函数
│       ├── context.py          # 执行上下文管理
│       └── loader.py           # 模型权重加载
└── assets/                     # 资源文件

workflow

时间轴:
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

初始化:
  [Rank 0] ModelRunner ✓ (CUDA Graph录制)
  [Rank 1+] ModelRunner ✓ (若tensor_parallel_size > 1)

Step 1(Prefill - Req1):
  Scheduler: 从 waiting 取 Req1
  Model:    处理 [token_ids_req1] (长度: 100)
  Sampler:  生成第101个token
  Update:   Req1 进入 running

Step 2(Prefill - Req2):
  Scheduler: 从 waiting 取 Req2
  Model:    处理 [token_ids_req2] (长度: 50)
  Sampler:  生成第51个token
  Update:   Req2 进入 running

Step 3(Decode - Req1, Req2):
  Scheduler: 从 running 取 [Req1, Req2]
  Model:    处理 [last_token_req1, last_token_req2] (每个batch 1个)
  Sampler:  采样 2 个 token
  Update:   Req1已生成56/256, Req2已生成2/256
  
Step 4(Decode - 继续):
  ......循环直到所有请求达到max_tokens或EOS

完成:
  返回: [{"text": "...", "token_ids": [...]}, {"text": "...", ...}]

Star History

Star History Chart

About

Nano vLLM

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages