← r/LocalLLaMA
▲
24
 
9👁
r/LocalLLaMA · u/bakawolf123 · 14d ago

PSA for M5Ultra owners running LLMs: set your prefill step to 8192

Prefill step size affects both PP performance and your drafter fetching logits for MTP (affects dflash as well, mlx-vlm needs to be patched slightly to support chunked prefill for dflash). It also needs to be reasonably large to be able to fill all your cores but not overfill - otherwise it will require more dispatches. In my tests I observe large gains up to 8k, e.g.: GLM-flash-4bit with MTP --prefill-step-size 8192 on raw mlx-vlm: Trial 1 (32768 prompt tokens): prompt_tps=1056.033, generation_tps=72.722, total_time=38.082 Trial 2 (65536 prompt tokens): prompt_tps=919.958, generation_tps=73.671, total_time=78.203 Trial 3 (131072 prompt tokens): prompt_tps=735.545, generation_tps=71.067, total_time=185.435 GLM-flash-4bit with MTP --prefill-step-size 2048: Trial 1 (32768 prompt tokens): prompt_tps=860.489, generation_tps=50.011, total_time=48.339 Trial 2 (65536 prompt tokens): prompt_tps=785.604, generation_tps=51.245, total_time=93.425 Trial 3 (131072 prompt tokens): prompt_tps=623.588, generation_tps=50.843, total_time=220.288 omlx with MTP (total time is skewed as it's 128TG vs 512 above): pp32768/tg128 44136.5 17.19 742.4 tok/s 58.6 tok/s 46.353s 709.7 tok/s 176.66 GB pp65536/tg128 86749.1 21.21 755.5 tok/s 47.5 tok/s 89.504s 733.6 tok/s 177.15 GB pp131072/tg128 178622.6 19.01 733.8 tok/s 53.0 tok/s 181.156s 724.2 tok/s 178.45 GB note that some engines (like omlx) support adaptive step size, e.g. the prefill speeds I observed for qwen3.8-flash-next on omlx even though it's starting from 2048 matches 8k performance from raw mlx-vlm at 64k context and above and even works 10% better on smaller context, but as you can see it's not always the case.

26 0 24 10/3 06:32 10/4 11:45 UTC
scorecomments9 sightings
first seen 2026-10-03 06:32 UTClast seen 2026-10-04 11:45 UTCscore then 24score now 24gained 0sightings 9
open on reddit ↗ 💬 17