If you are running Qwen 3.8 Flash Next on Strix Halo, use this software for inference. It's so much faster than llama.cpp especially at high context.
Here's the project. I have nothing to do with it. I'm just an amazed user.
https://github.com/gufo-org/gufo/blob/main/docs/models/qwen3.8-flash-next/BEN…
Those benchmark numbers hold up on real work loads. Here are some numbers I got during a chat.
"[6204 chunks in 119.0 s | encode: 1239 tok/s | decode: 57 tok/s]"
That's with MTP on. The PP speed in particular is just so fast. That PP speed is twice the speed of the fastest Strix Halo specific fork of llama.cpp I've ever used. Needless to say, the uplift is even greater compared to mainline llama.cpp.
It works with models other than QFN, but the current number is small. You can find the list on their project page.
scorecomments23 sightings
first seen 2026-10-03 06:30 UTClast seen 2026-10-07 08:02 UTCscore then 30score now 25gained -5sightings 23
open on reddit ↗
💬 49 (+3)