Highlights
- Pro
Stars
Docker configuration for running VLLM on dual DGX Sparks
GLM-5.2 QuantTrio TP=4+DCP2 on 4x NVIDIA DGX Spark (GB10) — tuned and measured under a real agent workload. Adaptive MTP, the indexer law, and the negative results.
DeepSeek-v4-Flash 0731 recipe for 2x DGX Sparks
One MiniMax H3 FL2VA video generated cooperatively across two NVIDIA DGX Sparks.
My implementation of Graspnet Graspness.
Intercept any app, then call it from Python like a library
A Python DSL to write Nvidia PTX for Hopper and Blackwell in JAX and PyTorch
Instant, Concurrent, Secure & Lightweight Sandbox for AI Agents.
MiniMax-M3 (428B, no pruning) at 36 tok/s on 2× NVIDIA DGX Spark — W4A16 GPTQ + NVFP4 KV + EAGLE-3 speculative decoding on vLLM. Three serving lanes: speed / balanced / long-context.
A tiny ~10K-parameter LLM router that learns which open-source model (deepseek-v4-pro / glm-5p2 / kimi-k2p6 via Fireworks) should answer each question and in what role, trained by evolution (sep-CM…
Recipe: GLM-5.2 (unpruned QuantTrio Int4-Int8Mix) at 200K ctx with MTP spec decode on a 4x NVIDIA DGX Spark (GB10) cluster
A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downs…
drowzeys / Keys---Full-GLM-5.2-Quantrio-INT4-INT8-mixed-8bit-Attention-on-4-x-DGX-Spark-GB10-Cluster
Updated (Latest): Full (non-pruned) GLM-5.2 DSA serving TP=4 on 4×DGX-Spark/GB10 via rebuilt sm_121a vLLM — Updated now with NVFP4 KV + increased context to 100ktarter pack for 4 DGX-Spark (GB10) …
Official PyTorch implementation of Synergies Between Affordance and Geometry: 6-DoF Grasp Detection via Implicit Representations
Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
Z80-μLM is a 2-bit quantized language model small enough to run on an 8-bit Z80 processor. Train conversational models in Python, export them as CP/M .COM binaries, and chat with your vintage compu…
[NeurIPS 2025 Spotlight] Reasoning Environments for Reinforcement Learning with Verifiable Rewards
🧱 easy fast local-first microVM runtime and library
A Python module to bypass Cloudflare's anti-bot page.
A TTS model capable of generating ultra-realistic dialogue in one pass.