← r/LocalLLaMA
▲
28
+2
28👁
r/LocalLLaMA · u/jacek2023 · 7d ago

CUDA: fuse shared experts into MMVQ by am17an · Pull Request #29184 · ggml-org/llama.cpp

MoE speedup, but only for some MoE architectures (like Qwen 35B A3B)

33 0 28 10/3 06:28 10/7 19:50 UTC
scorecomments28 sightings
first seen 2026-10-03 06:28 UTClast seen 2026-10-07 19:50 UTCscore then 26score now 28gained +2sightings 28
open on reddit ↗ 💬 11 (+2)