Qwen 3.8 Flash Next q2_0 running on a 2060 laptop (32 GB RAM + 6 GB VRAM) using Strata!
OK, this engine is indeed the real deal. I have expected perhaps 30 token/s prompt processing and 3 token/s decode at max, because the full Qwen 3.8 Next has around 120B parameters (not counting Engrams) and since I have just 32 GB RAM I thought it would crawl to a halt with SSD swapping.
But 10 token/s at 50k context is simply amazing on such an old device and with such a large model! That figure really surprised me and is very usable in my opinion.
It's 4 bit kv cache and no vision, so comprimises have to be made. But for real, the prefill speed is the only thing that keeps this from being usable, almost 100 token/s prefill is much higher than I have anticipated, but you still wait a long while for it to process large prompts. Qwen A35b A3B has around 5x faster prefill, and allows me to use 100K context without having to quant the kv cache at all. So not quite a replacement for that, but who knows if more optimizations are coming?
In any case, this is a very impressive showing. Used ./START-HERE.bat --draft-vocab en --vram-reserve-mib 100 --kv q4\_0 to run it.
This engine really deserves the hype it gets.