Qwen 3.8 flash for a single spark
Hi all proud to release my low bit quant for the qwen 3.8 flash base.
It retains 95% of b16 accuracy and doesnt suffer degradation at longer contexts.
Im still working on further inference improvements but you will see \~40tok/s.
Importantly, this quant allows for headroom to run both 256kl and RoPE.
Please enjoy!
https://huggingface.co/DJLougen/Qwen3.8-Flash-Next-Mooney
scorecomments5 sightings
first seen 2026-10-03 06:28 UTClast seen 2026-10-07 05:39 UTCscore then 0score now 0gained 0sightings 5
open on reddit ↗
💬 7 (+4)