← r/LocalLLaMA
▲
105
+3
18👁
r/LocalLLaMA · u/feelspeaceman · 31d ago

For Strix Halo - Official llama.cpp isn't ideal and how to highest possible throughput

I've been making a lot of comments about optimal setup for Strix Halo (gfx1151) and from my observation, 90% of our community is using offcial llama.cpp for it, which is NOT optimized for Strix Halo at all, official llama.cpp is having extremely hard time to reach 50% hardware theory, wasting the silicon of this device.

Here's alternatives that can bring the speed of Strix Halo to a totally different world, I will link to user's sastifaction comment to prove that the result is real:

Note: Official llama.cpp running Qwen38FN at 2xt/s and 2xxt/s prefill - 50% theory.

Hopefully this will be helpful to the Strix Halo users.

106 0 105 10/3 04:45 10/7 17:06 UTC
scorecomments18 sightings
first seen 2026-10-03 04:45 UTClast seen 2026-10-07 17:06 UTCscore then 102score now 105gained +3sightings 18
open on reddit ↗ 💬 140