← r/LocalLLaMA
▲
240
+3
24👁
r/LocalLLaMA · u/Beamsters · 30d ago

Qwen3.8-Flash-Next on MLX-serve, 1m context is released!

post image

Hi, I'm the co-creator of this Qwen3.8-Flash-Next engine support in MLX-serve. I've been tuning this one to run both fast, efficient and correct up 1m context using kv cache 8 bits in M5 Max 128GB. Qwen is working well at very long context as showed in the video (a snapshot at \~760k context), I let it build MLX Serve Monitor plugin that you've seen on the right side of Opencode2's app. The generation sustain through 1m context at around 40 tok/s on prose and 75 tok/s on coding. This specific quant uses 8 bits for dense layers and 4 bits expert layers, so the model's quality remains very high.

I didn't build this engine to show off tok/s on a very short context, repetitive greedy generation, but a real temp 1.0 sampling through deep context work. You would need iogpu.wired\_limit\_mb=120000 before attempt 1mb full context, because it needs around \~117GB on peak memory usage. There will be bugs around here and there, I couldn't test everything and every single use case, pls report.

You can grab it here: https://github.com/ddalcu/mlx-serve
Model's weight: https://huggingface.co/ddalcu/Qwen3.8-Flash-Next-MLX-Serve-mixed-4-8bit
Opencode2's plugin: https://github.com/beamivalice/opencode2-mlx-serve

Launch parameters (for 1 concurrency)

--model ./llm/models/Qwen3.8-Flash-Next-MLX-Serve-mixed-4-8bit \
--host 127.0.0.1 \
--port 11234 \
--ctx-size 1048576 \
--kv-quant 8 \
--max-tokens 64000 \
--mtp \
--prefix-cache-mem 10GB \
--prefix-cache-entries 1 \
--ssm-checkpoint-max 16 \
--metrics

241 0 240 10/2 17:36 10/8 14:52 UTC
scorecomments24 sightings
first seen 2026-10-02 17:36 UTClast seen 2026-10-08 14:52 UTCscore then 237score now 240gained +3sightings 24
open on reddit ↗ 💬 63