Hand-written NVFP4 W4A16 CUDA kernels for Volta
-
Updated
Aug 19, 2026 - Python
Hand-written NVFP4 W4A16 CUDA kernels for Volta
The set-and-forget LLM engine for Pascal and Volta: PXQ codec + kernels, auto-tuned per card. Ready-to-run PXQ models in MODELS.md; benchmarks vs llama.cpp in the README.
NVIDIA Tesla V100 显卡驱动自动化管理套件
Benchmarks and notes for running modern LLMs with vLLM on 8x Tesla V100-32GB in 2026.
Qwen3.8-27B in native NVFP4/FP8 on 2x PCIe Tesla V100-32GB (SM70): the PCIe runbook for v100-skinny + 1Cat-vLLM, with the 3 fixes that make it work without NVLink. 61-74 tok/s decode, MTP speculative decoding, OpenAI-compatible.
Fork adding clock mode for GPUs without P-states (V100, P100, pre-Volta) — upstream PR #10. A daemon that automatically manages the performance states of NVIDIA GPUs.
Reproducible llama.cpp kernel and runtime optimization lab for dual NVIDIA Tesla V100 GPUs (SM70)
Multi-GPU acceleration for MiniMax H3 video generation on NVIDIA V100 (sm_70). Ulysses sequence parallelism as a drop-in ComfyUI custom node — ~19 min to ~7 min on 8x V100.
Running large LLMs on pre-Ampere NVIDIA hardware — Tesla V100 (sm_70), RTX 2080 Ti (sm_75), CMP 170HX. Measured benchmarks, vLLM forks, and the hardware side: NVLink on SXM2 carrier boards, driver traps, cooling, used-kit acceptance.
Serve Qwen3.8-Flash-Next (125B MoE, NVFP4) at 262K context on 4x Tesla V100-SXM2-32GB — a Volta (sm70) port of SGLang for agentic coding. Native OpenAI + Anthropic APIs.
SGLang fork for IBM POWER9 (ppc64le): Tesla V100 sm70, CUDA 12.4, Granite LLM inference. Triton attention, float16, OpenAI-compatible API.
Serve Qwen3.5-397B-A17B (AWQ) on 8x Tesla V100-SXM2-32GB (DGX-1, TP8) for agentic coding & ops — a downstream fork of 1Cat-vLLM.
Open hardware desktop AI node: 4× Tesla V100, 128GB HBM2, PCIe/NVLink topology and V-Core liquid/air cooling.
Runbook + benchmarks: Qwen3.8-Flash-Next-ABLITERATED NVFP4 on 4× Tesla V100-32GB (reflashed SXM2→PCIe, 2+2 NVLink + PLX). 1Cat-vLLM 1.5.0, TP4 — 262,144-token context validated, 46 tok/s decode, 122 tok/s aggregate at 4 concurrent streams. Full E0–E17 optimization log with measured evidence.
云途觉晓多人云音创作平台:基于 MOSS-TTSD 的简体中文多人语音、参考音色与 V100 部署方案
OpenAI Triton compiler fork for IBM POWER9/POWER10 (ppc64le) with CUDA 12.4. Tesla V100 sm70 GPU kernels for PyTorch and SGLang.
PyTorch 2.12 fork for IBM POWER9/POWER10 (ppc64le) with CUDA 12.4 and Triton. Tesla V100 sm70, GPU training and LLM inference.
Run Mistral Voxtral realtime speech-to-text on a Tesla V100 (sm_70/Volta) via a patched vLLM — RTF ~0.4, keeps up real time
To associate your repository with the tesla-v100 topic, visit your repo's landing page and select "manage topics."