DEV Community

#llamacpp

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
llama-server's prompt cache reuses KV computed under a different LoRA scale

llama-server's prompt cache reuses KV computed under a different LoRA scale

Comments
6 min read
Fine-tuning Qwen2.5-1.5B on a Mac to fix a parroting chatbot (and halve the prompt)

Fine-tuning Qwen2.5-1.5B on a Mac to fix a parroting chatbot (and halve the prompt)

Comments
6 min read
Running 100B+ MoE Models on a Single RTX 4090: A Practical Guide to Expert Offloading with llama.cpp

Running 100B+ MoE Models on a Single RTX 4090: A Practical Guide to Expert Offloading with llama.cpp

Comments
8 min read
Running 100B+ MoE LLMs on a Single RTX 4090: A Practical Guide to Expert Offloading with llama.cpp

Running 100B+ MoE LLMs on a Single RTX 4090: A Practical Guide to Expert Offloading with llama.cpp

Comments
7 min read
When Chat Templates Go Wrong

When Chat Templates Go Wrong

Comments
7 min read
Benchmarking Gemma 4 E2B on CPU with llama.cpp: A Practical Local AI Experiment

Benchmarking Gemma 4 E2B on CPU with llama.cpp: A Practical Local AI Experiment

Comments 4
10 min read
DFlash-2: Benchmarking Z-Lab's Successor to DFlash for Accuracy and Throughput Gains

DFlash-2: Benchmarking Z-Lab's Successor to DFlash for Accuracy and Throughput Gains

1
Comments 7
14 min read
llama.cpp vs Ollama in 2026: Which Runtime Should You Run?

llama.cpp vs Ollama in 2026: Which Runtime Should You Run?

Comments
18 min read
A 4 GB Laptop GPU vs a 6-Core CPU on Gemma 4, Re-Measured in ABBA Order: 4.1x

A 4 GB Laptop GPU vs a 6-Core CPU on Gemma 4, Re-Measured in ABBA Order: 4.1x

10
Comments 2
11 min read
ROCm vs Vulkan for AMD Local LLM Hosting: 2026 Guide

ROCm vs Vulkan for AMD Local LLM Hosting: 2026 Guide

Comments 2
20 min read
KV Cache on 16 GB GPUs: Making Long Context Actually Fit

KV Cache on 16 GB GPUs: Making Long Context Actually Fit

Comments 1
22 min read
How to pick --n-cpu-moe in llama.cpp: Qwen3.6 35B-A3B on 12, 16 and 24 GB GPUs

How to pick --n-cpu-moe in llama.cpp: Qwen3.6 35B-A3B on 12, 16 and 24 GB GPUs

1
Comments 2
3 min read
llama-server's logprobs are placeholders when speculative decoding is on

llama-server's logprobs are placeholders when speculative decoding is on

1
Comments 2
7 min read
Gemma 4 at Over 70 Tokens/s on a 2021 Laptop's 4 GB GPU: The Live Demo, Step by Step

Gemma 4 at Over 70 Tokens/s on a 2021 Laptop's 4 GB GPU: The Live Demo, Step by Step

9
Comments 2
12 min read
Inside My llama.cpp Setup: Tuning Qwen 3.8 27B for 512K Context

Inside My llama.cpp Setup: Tuning Qwen 3.8 27B for 512K Context

Comments 3
7 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.