llama : add a GPU cache for MoE experts kept in host memory by am17an · Pull Request #29887 · ggml-org/llama.cpp
Potentially big speedup for MoE models that don’t fully fit in VRAM.
Are you GPU Poor? Show your speedups ;)
update https://github.com/ggml-org/llama.cpp/pull/30112 MERGED
scorecomments24 sightings
first seen 2026-10-07 19:23 UTClast seen 2026-10-09 07:36 UTCscore then 119score now 436gained +317sightings 24
open on reddit ↗
💬 132 (+97)