mtp
Here are 246 public repositories matching this topic...
The fastest way to run Qwen 3.8 Flash Next and Qwen 3.8 27B on a Mac: 125 tok/s in OpenCode on an M5 Max. Native MTP speculative decoding on Apple Silicon, exact at any temperature. OpenAI and Anthropic compatible local server.
-
Updated
Sep 19, 2026 - Python
Android File Transfer for Linux (and macOS!)
-
Updated
Apr 17, 2026 - C++
This repository holds the source code of Microsoft.Testing.Platform (MTP), a lightweight alternative to VSTest, as well as MSTest adapter and framework.
-
Updated
Sep 23, 2026 - C#
A modern Android File Transfer tool for macOS with AI supercharged.
-
Updated
Sep 13, 2026 - Swift
llama.cpp fork with TurboQuant WHT-rotated KV cache & weight compression + Gemma 4 MTP and Qwen 3.6 NextN speculative decoding (+30-50% throughput).
-
Updated
Sep 22, 2026 - C++
Lightweight USB Media Transfer Protocol (MTP) responder daemon for GNU/Linux
-
Updated
Jun 5, 2026 - C
Qwen 3.8 27B ROCmFP4 on AMD Strix Halo (Ryzen AI Max+ 395). Up to 36 tok/s via MTP Speculation, TurboQuant & Mesa RADV Wave64.
-
Updated
Sep 7, 2026 - Python
A library to access MTP (Media Transfer Protocol) Devices.
-
Updated
Sep 22, 2026 - C
Open recipes, engine patches, and benchmark harnesses for LLM inference on Intel Arc Pro B60/B70 (Battlemage, Xe2). MoE 35B at 160 t/s decode / 7.5K t/s prefill single-stream, 27B at 50~ t/s decode / 1.7K t/s prefill single stream. vLLM XPU MTP unlocked. Muse Glimmer recipe added!!
-
Updated
Sep 23, 2026 - Python
llama.cpp fork for significantly improved performance on Ampere (especially RTX 3090 / 3090 Ti): TurboQuant KV cache, MTP speculative decoding with a 64K draft-vocabulary shortlist, custom SM86 + Qwen kernels. 90 tok/s over a 100K-token generation at temperature 1.
-
Updated
Sep 19, 2026 - C++
Qwen4-Exp (Qwen3.8-Flash-Next) SWA + MTP + PLE-on-iswa — bounded deep-context decode w/ long-range recall, TBQ4 4.25 bpv KV, DSpark + RotorQuant. Main (master) = Qwen4-Exp build; the Qwen35 SWA hybrid lives on the qwen35 branch. RTX 4090.
-
Updated
Sep 22, 2026 - C++
Qwen3.8-Flash-Next on 2× RTX 3090 + 128 GB RAM. Experimental peaks: 1,860 tok/s prefill at 260K input; 89.1 tok/s decode at 128K input. Full 256K context.
-
Updated
Sep 18, 2026 - Python
NeveAI é uma plataforma de IA local privacy-first, desenvolvida para oferecer uma experiência de alta performance na execução de LLMs, reduzindo a dependência de grandes plataformas, assinaturas caras e APIs externas e servindo como uma alternativa offline, privada e independente.
-
Updated
Sep 22, 2026 - Python
Add this topic to your repo
To associate your repository with the mtp topic, visit your repo's landing page and select "manage topics."