What's the current meta for RDNA4 with Qwen 3.8?
As it says, really - I'm currently running vllm-radiance on dual R9700s, with Qwen 3.8 27B FP8 (or, rather, Swift 1.5 FP8). Performance is great an' all (5000t/s prefill, 130t/s+ code gen), but I'm just wondering...with all the architecture-specific inference engines popping up all over the place...is there anything I'm missing out on? I couldn't find anything that could give better performance on RDNA4 when I looked, so...over to you guys?
I'm particularly interested in anything that could potentially get up and running with Qwen 3.8 Flash Next - vllm-radiance doesn't support it yet, but I don't particularly want to regress to the performance of llama.cpp after having experienced vllm-radiance performance levels.
scorecomments16 sightings
first seen 2026-10-06 09:47 UTClast seen 2026-10-09 08:00 UTCscore then 3score now 2gained -1sightings 16
open on reddit ↗
💬 23 (+17)