← r/LocalLLaMA
▲
11
-1
8👁
r/LocalLLaMA · u/nirurin · 13d ago

Best current Qwen Flash Next Q4-ish? + worth using?

Im running a 5090 and 64gb of ram, so im limited on what I can run. I have currently been able to fit the following - Atomic Q4\_k\_m 4.27bpw @ 31 layers offload Swift IQ4\_xs @ 32 layers offload. Im about to try the Unsloth IQ4\_xs as well. I could get a "bigger" (non IQ) quant for atomic because its smaller, however they do theirs is obviously different. The unsloth IQ4 is also pretty small, the Swift one is the biggest. i may be able to jump up one size on something, but it would mean offloading more layers and that would seem to be a significant slowdown. I get around 40tok/s if I stay around the 34-30 range. any recommendations? and the next question - I can (and do) also run Q5 and Q6 qwen 27b models. Is the bigger quant of 27b actually going to be more intelligent than the cut-down flash-next builds?

12 0 11 10/3 06:31 10/4 23:51 UTC
scorecomments8 sightings
first seen 2026-10-03 06:31 UTClast seen 2026-10-04 23:51 UTCscore then 12score now 11gained -1sightings 8
open on reddit ↗ 💬 21