Starred repositories
3
results
for forked starred repositories
Clear filter
llama.cpp fork with TurboQuant WHT-rotated KV cache & weight compression + Gemma 4 MTP and Qwen 3.6 NextN speculative decoding (+30-50% throughput).
TheTom / llama-cpp-turboquant
Forked from ggml-org/llama.cppLLM inference in C/C++
z-lab / llama.cpp-fork
Forked from ggml-org/llama.cppLLM inference in C/C++