← r/LocalLLaMA
▲
4
 
1👁
r/LocalLLaMA · u/zipzak · 3h ago

Qwen Flash Next @ 137 tok/s & 3,497 tok/s Prefill w/ 512k context on a 5090, 192gb ram, Windows Build, comparing Strata and Infernix

Strata vs Infernix on the same model, RTX 5090 — A/B at 262k and 512k context

Hi folks, another inference engine post for your feed. I was blown away by Strata, but saw someone else post about Infernix with some wild claims, and it doesn't seem so well known, so I've spent a day downloading models and A/B testing each engine. Completely stunned that either of these pieces of software run on Windows, and the setup was relatively painless too.

Same-day interleaved benches, same model (Qwen3.8-Flash-Next Uncensored, orcarouter's Apache-2.0 abliterated checkpoint), same inference contract (int8 KV, MTP spec-4 + lm-head-draft, YaRN 2 for 512k). 3 decode runs per cell, medians.

Hardware: RTX 5090 32 GB (Gen5 x16), 189.6 GB RAM, 9950x, Windows 11, models on NVMe, llama-swap fronting both.

Engines:

  • Strata 0.1.41, orcarouter Q4\_K\_S GGUF, calibrated flags (pcie-frac 0.55, pool-workers 4, adapt-every 1/160/0.97, spec-min-p 0.7)
  • Infernix 2026.10.09.2, published NVFP4 recipe-C artifact (76.3 GB) + shared 52 GB n-gram volume on NVMe (mmap'd, never loaded to RAM — the 63.3 GiB of experts is what's pinned). Vision works (--vision, tower offloaded to pinned RAM, borrows VRAM only while encoding)

Results (decode tok/s, median of 3 / cold prefill tok/s):

||262k|512k (YaRN 2)|
|:-|:-|:-|
|Strata Q4\_K\_S|120.6 / 1,550|126.3 / 3,045|
|Infernix NVFP4|149.0 / 3,548|137.0 / 3,497|

Takeaways:

  1. Infernix +23% decode at 262k, +9% at 512k — same model, same spec contract.
  2. 512k is nearly free (≤8% on either engine). On a 32 GB 5090 experts are the bottleneck, not attention KV.
  3. Prefix cache makes repeat long prompts free (\~0.09 s warm re-prefill of a 39k prompt).
  4. Quality: statistically tied — teacher-forced PPL 4.772 (Infernix) vs 4.844 (Strata's quant), ΔNLL −0.015 ± 0.010 nats; the artifact runs the abliterated weights bit-exact.
  5. Vision confirmed working on both engines — Infernix needs --vision (off by default) and hard-validates the model string (--model-id to match your proxy's entry name — got me with a 404).
  6. Cold start: Infernix \~28 s; Strata \~1 min.

Caveats: decode measured at \~39k effective context (deep-filled 500k is a different KV-vs-cache question); ±10-15 tok/s run noise, so treat sub-10% deltas as noise; YaRN past 262k is experimental (speed benched, not 512k quality).

Verdict: Infernix NVFP4 is the daily driver — \~150 tok/s, 512k, vision, all on one 5090. Strata stays as the second engine. Using it for coding and other tasks, it's fantastic at creative writing / role play in Silly Tavern, and I've been using it instead of smaller creative finetunes. For agentic work and tasks, it has been more than capable of handling everything I've asked of it, and has corrected mistakes from Deepseek 4.1 and GLM 5.3 Flash that I had to pay them to write over api.

posted Sat, 10 Oct 2026 00:14:28 GMTseen 1 time
open on reddit ↗ 💬 11