Become a sponsor to blackwellboy
I do independent behavioral and performance testing of local AI models on real hardware: a multi-node DGX Spark fleet running the models people actually self-host. The published work includes full tuning sweeps with every losing config, a 12-hour production soak with raw per-turn logs, container recipes with pinned digests, and cross-validation with the offlabel behavioral guide project. Nothing is a leaderboard read; every number ships with the condition it was measured under and the raw data to check it.
Sponsorship funds hardware time, longer soaks, and more models characterized. The results stay public, published the same way: raw logs, losing cells included, claims scoped to what was actually measured.
Featured work
-
Blackwellboy/laguna-s21-lab
Laguna S 2.1 (NVFP4) on DGX Spark GB10: 20-cell tuning sweep, 12h soak (3,099 turns, raw logs), thinking-gate dose-response study, Qwen 35B head-to-head, quant-floor reference. Losers published too.
Python 9 -
Blackwellboy/Hy3-295B-NVFP4-MTP-Dual-DGX-Spark
First public recipe for Tencent Hy3-295B (295B MoE, NVFP4) across two NVIDIA DGX Spark GB10 nodes: vLLM TP=2, MTP speculative decoding, and the full toolchain-debugging story including what fails.
Shell 2 -
Blackwellboy/MiniMax-M3-2x-DGX-Spark-stock-driver
MiniMax-M3 (428B) at 34-41 tok/s on 2x DGX Spark GB10 on the stock CUDA-13.0 driver (thinking-off, temp 0, single stream). KVarN + EAGLE-3 recipes, patches, prebuilt image, full runbook.