← r/LocalLLaMA
▲
1
 
10👁
r/LocalLLaMA · u/DimeRhyme · 11d ago

An UltraFast Qwen3.8 Flash recipe: 74 tok/s, 212 tok/s aggregate on one DGX Spark

post image

TL;DR: vLLM recipe for Qwen3.8 Flash on one DGX Spark / GB10. 74 tok/s peak single-stream, 60 to 70 on normal requests, 212 tok/s across 8 streams, 2x to 3.4x faster cold prefill than the recipe it's forked from, full 262K context, and quality matches the original within noise. Everything is open, including the raw per-round data and the benchmark scripts.

https://github.com/dime-online/qwen3.8-Flash-DGX-UltraFast

Most single-Spark setups I've seen posted for this model land somewhere in the 35 to 45 tok/s range, so I spent a few weeks figuring out where the time per token actually goes on a GB10 and cutting it down. If you don't have a Spark, the tricks in the middle section should still be interesting, since most of them apply to any MTP or speculative decoding setup.

What this is, in plain terms

It's a ready-made serving setup. You build the container, pull the public weights, and get an OpenAI-compatible server that answers a lot faster on one box. The speed comes from the model's own draft head guessing several tokens ahead while the full model checks all of them in one pass. Right guesses give you several tokens for the price of one step, and wrong ones get replaced by the full model's answer, so output quality doesn't change.

Decode

Peak decode speed, same workload at every point, best of 3 rounds:

|Streams|1|2|3|4|5|6|7|8|
|:-|:-|:-|:-|:-|:-|:-|:-|:-|
|tok/s|74.1|110.0|132.9|155.5|175.8|191.5|205.8|212.2|

Tokens per step stays between 3.65 and 3.94 from 1 all the way to 8 streams, so the speculation doesn't fall apart under batching. 8 is where it tops out because that's the configured max\_num\_seqs, and the gain from 7 to 8 was down to 3%.

Prefill

Cold prompt with nothing cached, three repeats each:

|Prompt|This recipe|Original recipe|Increase|
|:-|:-|:-|:-|
|16K tokens|4,016 tok/s|1,171 tok/s|\+243%|
|64K tokens|2,426 tok/s|1,071 tok/s|\+127%|
|128K tokens|2,213 tok/s|1,065 tok/s|\+108%|

With prefix caching on, a cached coding prompt starts replying in about 0.57 s, which is what makes agent loops feel fast.

What actually made the difference

The model's own MTP head, run densely, lands about 3.7 tokens per verify step. That's the single biggest lever.

I cut the draft head's vocab from 248K to 65K ids. On a GB10 the draft pass is memory-bound, and reading a full-vocab head every draft step was a real chunk of the step time. The target model still verifies against the full vocab, so this can only change speed, not output.

The quant is W4A16 AutoRound for the MoE experts, FP8 for the side layers and INT8 for the lm\_head. No 3-bit and no NVFP4, because I wanted the speed to come from the serving path and not from squeezing the weights harder.

There's a GB10-tuned low-latency GEMM for the small decode-time matmuls and a sort-free top-k in the verify step.

The prefill gain mostly comes from a faster gather path for the per-layer embedding table, which removes a pile of serial page faults during prefill.

Put together, each decode step went from 68.3 ms on the original recipe to 52.3 ms on an agent-shaped coding workload, about 1.3x more steps per second.

One thing that didn't pay off: doubling the prefill chunk to 16,384 tokens gave no prefill gain at all and ran the box low enough on memory that I rejected it.

Quality

93.1% and 93.3% on a fixed 492-question suite over two seeds, covering code with execution checks, math, knowledge, instruction following, tool calls and long-context needles. I also ran a teacher-forced check against the original checkpoint, and top-1 agreement moved by 0.06 points against a 0.15 point noise band I set before running it.

Practical stuff

The model takes about 71 GiB, the KV pool is 16 GB, and around 16 GiB stays free under load. While generating, the GPU draws about 35 to 37 W median and peaks near 70 W on long cold prefills, with no power or thermal throttling across the soak runs.

The 65K draft vocab was built from English and code, so Chinese, Japanese and Korean output drafts less well and runs slower. Quality isn't affected, because the full model still checks every token.

Built on Saren-Arterius's qwen3.8-Flash-DGX-AutoRound, so big credit there. Happy to answer questions.

2 0 1 10/3 06:30 10/7 08:01 UTC
scorecomments10 sightings
first seen 2026-10-03 06:30 UTClast seen 2026-10-07 08:01 UTCscore then 1score now 1gained 0sightings 10
open on reddit ↗ 💬 13 (+3)