low-bit
Here are 13 public repositories matching this topic...
Get more intelligence from every bit. Better quantization formats and smarter calibration let larger, stronger models run smoothly on the hardware you already own.
-
Updated
Sep 24, 2026 - C++
QuantFace: Towards Lightweight Face Recognition by Synthetic Data Low-bit Quantization
-
Updated
Jun 26, 2022 - Python
admm for cnn layerwise weight low bit quantization
-
Updated
Sep 25, 2019 - Python
Memory-efficient PyTorch optimizers with Adaptive Log-Space quantization, fused CUDA kernels, and Adafactor, CAME, APOLLO support.
-
Updated
Sep 9, 2026 - Python
2.898-BPW Qwen3-8B with direct-packed CPU/CUDA inference
-
Updated
Aug 7, 2026 - Python
Efficient low-bit KV-cache compression research with honest metadata accounting.
-
Updated
Apr 24, 2026 - Python
Pure-Julia CPU inference engine for BitNet b1.58 ternary LLMs
-
Updated
Jun 17, 2026 - Julia
Microscaling (MX) low-bit quantization formats, implemented and benchmarked on VLMs.
-
Updated
Aug 16, 2026 - Jupyter Notebook
Compute-for-memory research lab for running larger LLMs on memory-constrained consumer GPUs.
-
Updated
Sep 23, 2026
Quantization toolkit for large language models: run and store huge models on small VRAM. Inspired by Unsloth's dynamic quantization.
-
Updated
Sep 11, 2026
-
Updated
Aug 16, 2026 - Jupyter Notebook
Add this topic to your repo
To associate your repository with the low-bit topic, visit your repo's landing page and select "manage topics."