M5 Ultra - Qwen3.8 Flash Next vs Laguna S 2.1
Spent today running a same-day, same-harness shootout between Qwen3.8-Flash-Next (oMLX, 182GB oQ8e, MTP) and Laguna-S-2.1 GGUF (LM Studio, 128GB, 8bit) on a Mac Studio M5 Ultra 256GB. Both capped at 262K context, thinking on, unique content per run with zero cached tokens verified each time.
That last part matters because my first run was wrong hah... shared prefixes across sizes let the KV cache carry over and 200K "prefilled" in 21s.
Prompt Qwen Laguna
8K 2.0s 10.3s
32K 7.4s 30.4s
64K 14.7s 70.8s
131K 30.1s 217.2s
200K 47.1s 455.4s
Qwen holds \~4,200 tok/s linear which is amazing. Laguna degrades superlinearly (quadratic attention doing quadratic attention things). At 200K, prefill is 94% of total time on both.
Decode (tok/s): Qwen 59-74 across sizes (MTP at 70-76% acceptance per server logs, roughly 2x). Laguna 68 down to 34 as context grows. No speculation on Laguna, its DFlash path already lost to plain decode on this hardware in earlier testing. I think if Laguna could get DFlash figured out or MTP, this might be a different conversation.
Quality was a draw, 4/4 each, on four problems with script-verified answers (Muse created the gymnastics here: exact 9-digit combinatorics, interval code with 12 hidden tests, fresh knights/knaves, asyncio ordering trap). Opposite styles though: Laguna answers in 5-10s with a few hundred tokens, Qwen deliberates exhaustively (one answer took 119s / 11K tokens). Both burned a full 8K budget on hidden reasoning with zero visible output exactly once, then converted on a 16K retry.
Happy to answer methodology questions. Full writeup with charts and the test rig diagram: https://echalupa.com/blog/qwen-flash-next-vs-laguna-200k