← r/LocalLLaMA
▲
263
-3
25👁
r/LocalLLaMA · u/ilintar · 27d ago

Qwen3.8 Flash Next now at 1.2k t/s prefill on Strix Halo

As you all know, Qwen3.8 Flash Next on mainline llama.cpp is still in a pretty experimental stage, but a lot of community forks are trying to get it to work better. There's also a closed-source solution called Halogen (https://github.com/peonist-ai/halogen-flash-server) that boasted 1.2k t/s prefill numbers when the community fork barely reached 400. Since I dislike closed source and I like open source, I decided to take the challenge and bring llama.cpp up to the same performance level and I'm happy to report that after burning through a few evenings and a lot of tokens I got there.

The link contains the (Opus-generated) recap of the entire debugging / optimization journey made, as well as of course the links to the branch, the custom HIP runtime (making a glorious return), a probably-not-working-on-the-first-try installation script and the numbers. I'm trying to also explain the ecosystem - how the mainline, the community forks and custom forks/branches like mine work. Now that I've gotten the result, I'll work on cleaning it up and submitting proper PRs to mainline (will also submit a clean PR to the community fork), this should also help the GLM 5.3 Flash architecture since they use similar sparse attention.

270 0 263 10/2 17:36 10/9 00:51 UTC
scorecomments25 sightings
first seen 2026-10-02 17:36 UTClast seen 2026-10-09 00:51 UTCscore then 266score now 263gained -3sightings 25
open on reddit ↗ 💬 59