Advice needed on a budget hybrid build for Qwen3.8-Flash-Next at 4-bit
After seeing how good the cloud models are getting, I feel like this is something we cannot let big tech hold over us so deiced to build a budget local box.
After testing about 20 open models, Qwen3.8-Flash-Next (medium reasoning) was the only one that passed my task without inventing config options when used with a harness that forced doc lookups. So the box is built around that model. GLM-5.3-Flash performed even better but it's too big for my budget.
Planned build (Netherlands prices):
- Ryzen 5 9600, about €200
- MSI B850 Gaming Plus MAX WiFi, about €170
- 2×48 GB DDR5-5600, €1,199–1,549. Two sticks only to avoid the four-stick speed penalty.
- Used RTX 3090, about €1,150–1,500
- Case, 850 W PSU and NVMe I already own
That's about 79 GB of the model in RAM (experts plus the 28.8 GB n-gram table) and about 20.6 GB on the card.
Questions:
- Will 6 Zen 5 cores hold back generation with 40 MoE layers on the CPU?
- At 96 GB with about 79 GB of mode will 17 GB be enough for the OS, a sandbox container and an embedding model? Should I use
--mlock? - Is DDR5-6000 worth it over 5600?
- Has anyone run Unsloth's MTP branch with experts on the CPU? What speedup did you get, and does it break the prompt cache on the DeltaNet layers?
- Is anything wrong with a used 3090 here? Also has anyone tried the Arc Pro B60 (€772 new) workable on Vulkan or SYCL with this model yet? It's so much cheaper but I am worried becase of the software.
Super exciting to work on it but I am really inexperienced so this would be my first build. Does it make sense?
scorecomments12 sightings
first seen 2026-10-06 15:47 UTClast seen 2026-10-09 08:00 UTCscore then 5score now 5gained 0sightings 12
open on reddit ↗
💬 13 (+2)