← r/LocalLLaMA
▲
10
+8
14👁
r/LocalLLaMA · u/tabletuser_blogspot · 3d ago

MI50 ROCm 10.2 TheRock vs Vulkan Mesa 26.2 llama.cpp benchmarks

I prefer running llama.cpp Vulkan prebuilt binary. I just download the latest version and ready to roll. I finally took the hours necessary to get TheRock latest tarball version of ROCm 10.2 running on dual AMD Radeon Instinct MI50 gfx906 (32gb combined VRAM).

Same models benched in previous post. A mix of Dense and MoE models and quants that better utilize available VRAM.

  • llama_bench_Swift-Qwen3.8-27B-Uncensored-MTP.Q6_K.gguf
  • llama_bench_Swift-1.5-Qwen3.8-27B-Q6_K.gguf
  • llama_bench_Gemma-4-MoonGem-31B.i1-Q6_K.gguf
  • llama_bench_gemma-4-31B-it-UD-Q6_K_XL.gguf
  • llama_bench_Nemotron-3.5-30B-A3B-Antislop-FTPO.i1-Q5_K_M.gguf
  • llama_bench_Laguna-XS-2.1-APEX-I-Balanced.gguf
  • llama_bench_Agents-A1-Q4_K_M.gguf
  • llama_bench_Qwen3.6-35B-A3B-UD-Q5_K_XL.gguf
  • llama_bench_Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-Q6_K_P.gguf

Here is the backend performance comparison contrasting the native ROCm (v10.2 for gfx906) runtime against your optimized Vulkan (MESA\_PPA\_26.2) baseline. The data highlights a massive architectural split: ROCm significantly accelerates token generation across the board but suffers high variance and regressions in MoE pre-fills.

Both GPUs are power limited to 145 watts. Build versions used:

llama-b11382 used for Vulkan (prebuilt ubuntu binary)
llama-b11401 used for ROCm (compiled with proper flags)

Architectural Performance Breakdown: Vulkan vs. ROCm

|Model|Size|Params|Test|Vulkan Baseline (t/s)|ROCm 1st Run (t/s)|Performance Delta (%)|
|:-|:-|:-|:-|:-|:-|:-|
|qwen35 27B Q6\_K|20.88 GiB|27.32 B|pp512 tg128|149.30 ± 0.15 18.35 ± 0.01|176.86 ± 17.28 20.21 ± 0.65|\+18.46% +10.14%|
|qwen35 27B Q6\_K|22.21 GiB|27.32 B|pp512 tg128|163.48 ± 0.18 18.52 ± 0.03|183.49 ± 15.75 20.04 ± 0.63|\+12.24% +8.21%|
|gemma4 31B Q6\_K|23.46 GiB|30.70 B|pp512 tg128|119.98 ± 0.12 15.72 ± 0.02|172.93 ± 0.89 17.26 ± 0.13|\+44.13% +9.80%|
|gemma4 31B Q6\_K|25.62 GiB|30.70 B|pp512 tg128|133.67 ± 0.24 16.38 ± 0.02|178.59 ± 2.31 17.50 ± 0.09|\+33.61% +6.84%|
|nemotron\_h\_moe 31B|25.18 GiB|32.91 B|pp512 tg128|834.87 ± 1.72 63.82 ± 0.07|736.98 ± 134.77 103.86 ± 0.50|\-11.73% +62.74%|
|laguna 30B.A3B Q5\_K|22.64 GiB|33.44 B|pp512 tg128|723.34 ± 2.99 58.98 ± 0.28|856.14 ± 29.71 76.81 ± 0.20|\+18.36% +30.23%|
|qwen35moe 35B Q4\_K|19.70 GiB|34.66 B|pp512 tg128|957.62 ± 3.88 51.50 ± 0.07|834.58 ± 102.56 69.70 ± 0.22|\-12.85% +35.34%|
|qwen35moe 35B Q5\_K|24.76 GiB|34.66 B|pp512 tg128|909.20 ± 5.71 53.18 ± 0.07|763.93 ± 106.19 68.47 ± 0.34|\-15.98% +28.75%|
|qwen35moe 35B Q6\_K|28.53 GiB|34.66 B|pp512 tg128|763.05 ± 66.01 52.81 ± 0.04|754.82 ± 70.48 67.29 ± 0.21|\-1.08% +27.42%|

Core Insight Strategy & Bottlenecks

  1. Token Generation (tg128) Dominance: ROCm dominates pure text generation. The native AMD matrix kernels unleash your MI50 computation potential, unlocking a massive +62.74% boost for Nemotron but taking a small hit on pre-fill -11.73%.
  2. Dense Model Pre-fills (pp512): Dense architectures scale cleanly under ROCm. Gemma 4 sees a +33% to +44% processing throughput spike over the Vulkan RADV driver driver bounds. MoE models take a hit with a -15.89% difference with llama\_bench\_Qwen3.6-35B-A3B-UD-Q5\_K\_XL.gguf being the happiest with Vulkan backend.
  3. The MoE Prompt Processing Delinquency: Notice the massive standard deviations under ROCm for MoE pre-fills (e.g., Qwen3.5MoE Q4\_K has an instability block of ± 102.56). ROCm suffers from severe scheduling thrashing when building prompt streams across multiple active experts.

Here is the structured layout with pp512 and tg128 separated into individual columns for a clean side-by-side comparison between the two backend architectures.

Backend Comparison Table (Vulkan vs. ROCm)

|Model|Size|Params|Vulkan pp512 (t/s)|ROCm pp512 (t/s)|Vulkan tg128 (t/s)|ROCm tg128 (t/s)|
|:-|:-|:-|:-|:-|:-|:-|
|qwen35 27B Q6\_K|20.88 GiB|27.32 B|149.30 ± 0.15|176.86 ± 17.28|18.35 ± 0.01|20.21 ± 0.65|
|qwen35 27B Q6\_K|22.21 GiB|27.32 B|163.48 ± 0.18|183.49 ± 15.75|18.52 ± 0.03|20.04 ± 0.63|
|gemma4 31B Q6\_K|23.46 GiB|30.70 B|119.98 ± 0.12|172.93 ± 0.89|15.72 ± 0.02|17.26 ± 0.13|
|gemma4 31B Q6\_K|25.62 GiB|30.70 B|133.67 ± 0.24|178.59 ± 2.31|16.38 ± 0.02|17.50 ± 0.09|
|nemotron\_h\_moe 31B|25.18 GiB|32.91 B|834.87 ± 1.72|736.98 ± 134.77|63.82 ± 0.07|103.86 ± 0.50|
|laguna 30B.A3B Q5\_K|22.64 GiB|33.44 B|723.34 ± 2.99|856.14 ± 29.71|58.98 ± 0.28|76.81 ± 0.20|
|qwen35moe 35B Q4\_K|19.70 GiB|34.66 B|957.62 ± 3.88|834.58 ± 102.56|51.50 ± 0.07|69.70 ± 0.22|
|qwen35moe 35B Q5\_K|24.76 GiB|34.66 B|909.20 ± 5.71|763.93 ± 106.19|53.18 ± 0.07|68.47 ± 0.34|
|qwen35moe 35B Q6\_K|28.53 GiB|34.66 B|763.05 ± 66.01|754.82 ± 70.48|52.81 ± 0.04|67.29 ± 0.21|

So yes it's worth the hassle of jumping through hoops to get ROCm working on MI50 setups. At least I have 2 backends working. Next up I'll try some RPC.

https://preview.redd.it/ov4f9qmi2rth1.png?width=731&format=png&auto=w…

10 0 10 10/6 01:44 10/8 19:51 UTC
scorecomments14 sightings
first seen 2026-10-06 01:44 UTClast seen 2026-10-08 19:51 UTCscore then 2score now 10gained +8sightings 14
open on reddit ↗ 💬 4 (+4)