Should I not use MTP draft for agentic work?
I run qwen3.8-27b iq3\_s with llama.cpp to serve local hermes agent, in 16gb vram.
I noticed that enable MTP draft make prefill slower and the model seems less smart.
and vram is very tight I need to set the context length to 96k. decode speed can go around 40 to 60 tps.
if I disable MTP I can use 128k context but decode speed drop to like 35 tps.
What will you choose?
scorecomments17 sightings
first seen 2026-10-03 06:28 UTClast seen 2026-10-07 09:43 UTCscore then 1score now 0gained -1sightings 17
open on reddit ↗
💬 18 (+9)