← r/LocalLLaMA
▲
13
+6
8👁
r/LocalLLaMA · u/KingCpzombie · 22h ago

Best current R9700 inference engine?

There are way too many forks to keep track of, so I've gotten lost. As far as I can tell, Radiance VLLM is best for models that fit in GPUs while some form of llama.cpp is probably best for MOE RAM-spill?

My specific current goal is to run GLM5.3-Flash over 6 R9700s + system RAM but also looking to try Q-FN / DSv4-vision (or any other big models that I can fit, so not DSv4.1)

13 0 13 10/8 11:29 10/9 05:35 UTC
scorecomments8 sightings
first seen 2026-10-08 11:29 UTClast seen 2026-10-09 05:35 UTCscore then 7score now 13gained +6sightings 8
open on reddit ↗ 💬 19 (+17)