← r/LocalLLaMA
▲
365
+17
50👁
r/LocalLLaMA · u/KnownAd4832 · 15d ago

Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second

A while ago I posted 15 tok/s output and 100-120 tok/s prompt processing with the IQ3\_XXS quant on a 12GB RTX 5070 using llama.cpp. Since then I built my own inference engine for this one model and this kind of PC. The same IQ3\_XXS now runs at \~65 tok/s output and \~430 tok/s prompt processing, and the 2-bit quants run faster still using RCO-GSQ quantization.

Using:

64GB DDR5 (5600)
12GB RTX 5070 SFF (Gigabyte)
Ryzen 5 7600 CPU
Windows

Output (tokens/s) on 128K context:

Q2\_0 (equivalent to unsloth Q3): 65.1
IQ2\_XS (equivalent to unsloth Q4): 52.0
IQ3\_XXS (equivalent to unsloth Q5): 44.8

Prompt processing (tokens/s) on 128K:

Q2\_0: 543
IQ2\_XS: 472
IQ3\_XXS: 414

Requirements:

Q2\_0 = 37.6GB minimum in RAM+VRAM

IQ2\_XS = 39.2GB minimum in RAM+VRAM

IQ3\_XXS = 47GB minimum in RAM+VRAM

Vision encoder = 0.91GB additionally

You can now one click install and run the engine with low cost hardware (currently only optimized for CUDA).

GitHub: https://github.com/Niko1221/Strata

Model: https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF

366 0 365 10/2 17:29 10/9 02:52 UTC
scorecomments50 sightings
first seen 2026-10-02 17:29 UTClast seen 2026-10-09 02:52 UTCscore then 348score now 365gained +17sightings 50
open on reddit ↗ 💬 339 (+5)