← r/LocalLLaMA
▲
0
 
5👁
r/LocalLLaMA · u/Crampappydime · 7d ago

Qwen 3.8 flash for a single spark

Hi all proud to release my low bit quant for the qwen 3.8 flash base.
It retains 95% of b16 accuracy and doesnt suffer degradation at longer contexts.
Im still working on further inference improvements but you will see \~40tok/s.

Importantly, this quant allows for headroom to run both 256kl and RoPE.

Please enjoy!
https://huggingface.co/DJLougen/Qwen3.8-Flash-Next-Mooney

1 0 0 10/3 06:28 10/7 05:39 UTC
scorecomments5 sightings
first seen 2026-10-03 06:28 UTClast seen 2026-10-07 05:39 UTCscore then 0score now 0gained 0sightings 5
open on reddit ↗ 💬 7 (+4)