metal : few-row MMA mat-mul and batched copies for speculative decoding by pratiknarola-t · Pull Request #29869 · ggml-org/llama.cpp
Apple folks, it's for you.
llama-server with a Qwen3.8-27B DFlash2 Q8\_0 drafter, -ngl 99 -fa on -c 8192 -np 1 --jinja, DFlash2 with --spec-type draft-dflash --spec-draft-n-max 7. 64 generated tokens, median of 5 requests after one warm-up, mean of two server runs. Decode tok/s:
|mode|prompt|T|master|this PR|
|:-|:-|:-|:-|:-|
|serial|code|0|32.1|32.0|
|serial|code|1|32.1|32.0|
|serial|prose|0|32.1|32.0|
|serial|prose|1|32.1|32.0|
|DFlash2|code|0|30.2|110.0|
|DFlash2|code|1|24.3|80.9|
|DFlash2|prose|0|16.8|62.6|
|DFlash2|prose|1|13.9|48.8|
scorecomments8 sightings
first seen 2026-10-06 07:45 UTClast seen 2026-10-09 08:00 UTCscore then 3score now 3gained 0sightings 8
open on reddit ↗
💬 3 (+3)