CUDA: fuse shared experts into MMVQ by am17an · Pull Request #29184 · ggml-org/llama.cpp
MoE speedup, but only for some MoE architectures (like Qwen 35B A3B)
scorecomments28 sightings
first seen 2026-10-03 06:28 UTClast seen 2026-10-07 19:50 UTCscore then 26score now 28gained +2sightings 28
open on reddit ↗
💬 11 (+2)