← r/LocalLLaMA
▲
7
+5
18👁
r/LocalLLaMA · u/SomeITGuyLA · 6d ago

Qwen Flash next on 64GB RAM unified iGPU anyone ? (non-mac)

I've seen people reporting running it with 12GB VRAM + 64 GB RAM. Also with 64 GB RAM unified in Macs, but I was wondering if it's possible with any inference backend to run it for example on a 64 GB RAM minipc+ iGPU (780m in my case).
I'm currently running 125B Ling 3.0 flash at Q2 quants with llama.cpp (vulkan), its relatively usable, so I was wondering if a similar quant of Qwen Flash Next with the ngrams offloaded to SSD could work (even at low token/s). As far as I know this can't be done with llama.cpp now. Other inference engines does not seem to work with vulkan.

EDIT: Thanks everyone! It's working with the Q2 quant Qwen3.8-Flash-Next-GSQ-RCO-GGUF using llama.cpp with -lm mmap --lazy-mode on

8 0 7 10/3 15:28 10/7 23:52 UTC
scorecomments18 sightings
first seen 2026-10-03 15:28 UTClast seen 2026-10-07 23:52 UTCscore then 2score now 7gained +5sightings 18
open on reddit ↗ 💬 20 (+13)