← r/LocalLLaMA
▲
3
+2
8👁
r/LocalLLaMA · u/sdfprwggv · 4d ago

~188k warm ~60–67 tok/s: Qwen3.8-Flash-Next NVFP4 with Strata on a single RTX PRO 4500 32GB + 64GB DDR5.

Stack

  • Strata NVFP4 fork: github.com/sergqwer/strata-nvfp4
  • Model: jpezzulli/OrcaRouter-Qwen3.8-Flash-Next-Uncensored-ModelOpt-NVFP4 Hugging Face model
  • NVFP4 routed experts (\~63.3 GiB), separate FP8 PLE, INT8 KV cache, MTP speculative decoding
  • W4A8 prefill on Blackwell

Main engine flags:

./build/strata \
--pack packs/orca-nvfp4 \
--native models/orca-nvfp4.gguf \
--native-dense-gguf models/orca-nvfp4.gguf \
--ple-gguf models/ple-fp8.gguf \
--mtp mtp-orca/rt \
--spec 4 --spec-min-p 0.5 \
--prefill auto \
--expert-profile data/expert-profile.bin \
--expert-cache auto \
--resident-budget-gib 40 \
--max-context 200000 \
--kv int8

I serve it through Strata's OpenAI-compatible server.

Results so far:

  • \~50k context: up to \~80 tok/s
  • \~188k warm context: \~60–67 tok/s
  • cold 189k full prompt: \~1,680 tok/s prefill, \~53 tok/s decode

Pretty impressive for a single 32GB GPU + only 64GB system RAM.Running Qwen3.8-Flash-Next NVFP4 with Strata on a single RTX PRO 4500 32GB + 64GB DDR5.StackStrata NVFP4 fork: github.com/sergqwer/strata-nvfp4

Model: jpezzulli/OrcaRouter-Qwen3.8-Flash-Next-Uncensored-ModelOpt-NVFP4 Hugging Face model

NVFP4 routed experts (\~63.3 GiB), separate FP8 PLE, INT8 KV cache, MTP speculative decoding

W4A8 prefill on BlackwellMain engine flags:./build/strata \\
\--pack packs/orca-nvfp4 \\
\--native models/orca-nvfp4.gguf \\
\--native-dense-gguf models/orca-nvfp4.gguf \\
\--ple-gguf models/ple-fp8.gguf \\
\--mtp mtp-orca/rt \\
\--spec 4 --spec-min-p 0.5 \\
\--prefill auto \\
\--expert-profile data/expert-profile.bin \\
\--expert-cache auto \\
\--resident-budget-gib 40 \\
\--max-context 200000 \\
\--kv int8I serve it through Strata's OpenAI-compatible server.Results so far:\~50k context: up to \~80 tok/s

\~188k warm context: \~60–67 tok/s

cold 189k full prompt: \~1,680 tok/s prefill, \~53 tok/s decodePretty impressive for a single 32GB GPU + only 64GB system RAM.

3 0 3 10/5 19:39 10/7 11:27 UTC
scorecomments8 sightings
first seen 2026-10-05 19:39 UTClast seen 2026-10-07 11:27 UTCscore then 1score now 3gained +2sightings 8
open on reddit ↗ 💬 4 (+2)