The set-and-forget LLM engine for Pascal and Volta: PXQ codec + kernels, auto-tuned per card. Ready-to-run PXQ models in MODELS.md; benchmarks vs llama.cpp in the README.
pascal models cuda quantization volta p100 v100 pxq tesla-p100 llama-cpp vllm gguf speculative-decoding tesla-v100
-
Updated
Sep 22, 2026 - C++