Rx6800/rx6800xt gfx1030 and qwen 3.8 27b performance questions
Hi all,
After getting stuck in windows 11 llama.cpp and Vulkan, bugs and limitations on dual gpus, I moved to Linux and ROCm.
I just started but basically in windows 11/Vulcan, qwen3.8 27b unsloth q6\_k and ctk ctv at q8.0
\- sm layer with mtp on 35tok/sec TG (low context) and 180tok PP (due to a bug that cuts PP in half...)
\- sm layer without MTP 20tok/sec TG and 360 tok/sec PP.
\- sm tensor no mtp I get 15tok/sec TG 350 tok/s PP
Noticed better PP with small ub at 256
In Linux with ROCm no more MTP PP bug
\-sm tensor mtp on I get 45tok/sec TG and 450tok/sec PP.
So big progress but I have no idea how for far or close to performance ceiling of my GPUs.
Any new inference engine I should try?
I have not played with UB yet any other parameters to test?
Any numbers from other user on a dual gfx1030 to see PP TG numbers you guys get?
Thanks in advance!