Halogen + Qwen Flash Next keeps getting better
With latest Halogen version update (0.17.2), decode is consistently at \~45 tps even at high context with Qwen 3.8 Flash Next on a 128GB Strix Halo. This is some great work u/peonist-ai. Have been pumping out commit after commit with QFN. Its crazy good for a 177ish billion model. I dont think we are apprciating it enough 😂 Opus 5.5 plan implemented and reviewed by QFN is such high quality ❤️
https://preview.redd.it/yyji811t28uh1.png?width=1358&format=png&auto=…
scorecomments9 sightings
first seen 2026-10-08 11:29 UTClast seen 2026-10-09 07:36 UTCscore then 4score now 31gained +27sightings 9
open on reddit ↗
💬 56 (+53)