← r/LocalLLaMA
▲
436
+317
24👁
r/LocalLLaMA · u/jacek2023 · 40h ago

llama : add a GPU cache for MoE experts kept in host memory by am17an · Pull Request #29887 · ggml-org/llama.cpp

Potentially big speedup for MoE models that don’t fully fit in VRAM.

Are you GPU Poor? Show your speedups ;)

update https://github.com/ggml-org/llama.cpp/pull/30112 MERGED

437 0 436 10/7 19:23 10/9 07:36 UTC
scorecomments24 sightings
first seen 2026-10-07 19:23 UTClast seen 2026-10-09 07:36 UTCscore then 119score now 436gained +317sightings 24
open on reddit ↗ 💬 132 (+97)