← r/LocalLLaMA
▲
2
-1
16👁
r/LocalLLaMA · u/N34257 · 3d ago

What's the current meta for RDNA4 with Qwen 3.8?

As it says, really - I'm currently running vllm-radiance on dual R9700s, with Qwen 3.8 27B FP8 (or, rather, Swift 1.5 FP8). Performance is great an' all (5000t/s prefill, 130t/s+ code gen), but I'm just wondering...with all the architecture-specific inference engines popping up all over the place...is there anything I'm missing out on? I couldn't find anything that could give better performance on RDNA4 when I looked, so...over to you guys?

I'm particularly interested in anything that could potentially get up and running with Qwen 3.8 Flash Next - vllm-radiance doesn't support it yet, but I don't particularly want to regress to the performance of llama.cpp after having experienced vllm-radiance performance levels.

5 0 2 10/6 09:47 10/9 08:00 UTC
scorecomments16 sightings
first seen 2026-10-06 09:47 UTClast seen 2026-10-09 08:00 UTCscore then 3score now 2gained -1sightings 16
open on reddit ↗ 💬 23 (+17)