Skip to content
#

mtp

Here are 246 public repositories matching this topic...

The fastest way to run Qwen 3.8 Flash Next and Qwen 3.8 27B on a Mac: 125 tok/s in OpenCode on an M5 Max. Native MTP speculative decoding on Apple Silicon, exact at any temperature. OpenAI and Anthropic compatible local server.

  • Updated Sep 19, 2026
  • Python

Open recipes, engine patches, and benchmark harnesses for LLM inference on Intel Arc Pro B60/B70 (Battlemage, Xe2). MoE 35B at 160 t/s decode / 7.5K t/s prefill single-stream, 27B at 50~ t/s decode / 1.7K t/s prefill single stream. vLLM XPU MTP unlocked. Muse Glimmer recipe added!!

  • Updated Sep 23, 2026
  • Python

llama.cpp fork for significantly improved performance on Ampere (especially RTX 3090 / 3090 Ti): TurboQuant KV cache, MTP speculative decoding with a 64K draft-vocabulary shortlist, custom SM86 + Qwen kernels. 90 tok/s over a 100K-token generation at temperature 1.

  • Updated Sep 19, 2026
  • C++

NeveAI é uma plataforma de IA local privacy-first, desenvolvida para oferecer uma experiência de alta performance na execução de LLMs, reduzindo a dependência de grandes plataformas, assinaturas caras e APIs externas e servindo como uma alternativa offline, privada e independente.

  • Updated Sep 22, 2026
  • Python

Add this topic to your repo

To associate your repository with the mtp topic, visit your repo's landing page and select "manage topics."

Learn more