← r/LocalLLaMA
▲
0
-1
17👁
r/LocalLLaMA · u/Remarkable_Air_8383 · 6d ago

Should I not use MTP draft for agentic work?

I run qwen3.8-27b iq3\_s with llama.cpp to serve local hermes agent, in 16gb vram.

I noticed that enable MTP draft make prefill slower and the model seems less smart.

and vram is very tight I need to set the context length to 96k. decode speed can go around 40 to 60 tps.

if I disable MTP I can use 128k context but decode speed drop to like 35 tps.

What will you choose?

2 0 0 10/3 06:28 10/7 09:43 UTC
scorecomments17 sightings
first seen 2026-10-03 06:28 UTClast seen 2026-10-07 09:43 UTCscore then 1score now 0gained -1sightings 17
open on reddit ↗ 💬 18 (+9)