Qwen3.8-Flash-Next (5.05bpw + ngram at bf16) exl3 on one r9700: 863 t/s prefill and 35 t/s decode at 230k context (256k max), is that ok or am i missing something?
Hi, i got QFN running on my single r9700 but im not sure if i did everything right to get the best quality and speed out of this setup. Dont want to annoy anybody, maybe someone with the same card can tell me if this looks normal.
What i run:
\- model: Qwen3.8-Flash-Next from turboderp, exl3 5.05 bpw (head 6 bit, vision 6 bit, mtp 5 bit)
\- backend: exllamav3 rocm fork from phoenixhaxor (commit cbbef08), i had to patch one file for gfx12
\- cpu/ram: Ryzen 9 7945HX3D with about 90gb ram
\- 112 of 512 experts per layer are on the gpu, the other 400 on cpu with 16 threads
\- 262144 context, q8 kv cache, chunk size 4096, batch 1
\- mtp drafting is on, acceptance is around 53-54%
\- the ngram table (102gb, bf16 not quantized) gets streamed from nvme
Speed at 230k context (prose): prefill 863 t/s (only the new 101k tokens, the rest came from the prefix cache) and decode 34.7 t/s.
Is this ok for the r9700 or can i still tune something? Thanks 🙂