MTP in llama.cpp now decodes competitively with ds4 using GLM 5.3 Flash
Fine, pp is still slower, but I'm slowly coming around to the idea of MTP finally being useful on Apple Silicon, and this is the first time I'm seeing a model outperform ds4 (and that's with IngeniousIdiocy's M3U tuning). MTP seems to have no advantage there as was always the case with llama.cpp, until now it seems.
Qwen38FN will be the real test: vanilla ds4 currently spludging out 65 t/s (75 concurrently)...
Edit: I had no idea that MTP had such a massive impact on quality.... UNUSABLE and too bad...
scorecomments7 sightings
first seen 2026-10-08 19:30 UTClast seen 2026-10-09 07:36 UTCscore then 10score now 6gained -4sightings 7
open on reddit ↗
💬 9 (+5)