← r/LocalLLaMA
▲
8
 
1👁
r/LocalLLaMA · u/hyudryu · 3h ago

Strata with Qwen3.8 Flash Next UD-Q4_K_XL

Been seeing quite a few posts about Strata lately, so I figured I'd give it a shot on my RTX PRO 6000 and I am very impressed. Most of the benchmarks I've seen are using IQ2 or IQ3 quants, so I wanted to see how Unsloth's UD-Q4_K_XL performs instead. Setup: GPU: 1x RTX PRO 6000 Workstation (96GB) Model: Qwen3.8-Flash-Next UD-Q4_K_XL (Unsloth) Backend: Strata KV cache: INT8, 256K context Speculative decoding: MTP-4 Output: ~256 tokens per task 2 runs per task Single-request decode (C1): Task Strata (RTX PRO 6000) vLLM (RTX PRO 6000) vLLM (DGX Spark TP2) Prose 185.8 95.3 (-49%) 38.2 (-79%) Counting 320.4 171.0 (-47%) 75.3 (-77%) Coding 298.9 162.6 (-46%) 60.1 (-80%) Reasoning 289.1 149.6 (-48%) 51.8 (-82%) Both the VLLM instances were running Nvidia's NVFP4 quant so it's not exactly apples to apples, but the improved speeds are obvious. Qwen 3.8 flash next is flying through coding tasks and it's crazy how efficient it is. Looking forward to Qwen 4!

posted Sat, 10 Oct 2026 09:50:37 GMTseen 1 time
open on reddit ↗ 💬 12