← r/LocalLLaMA
▲
0
 
1👁
r/LocalLLaMA · u/Training_Visual6159 · 6d ago

Imma just say it, Strata absolutely clowned llama.cpp

So, I've been begging llama.cpp to do MoE caching for about a year, and watching them d ck around with 1% here and 2% improvements there instead... Until Strata (https://github.com/Niko1221/Strata) clowned llama with 5-10x prefill and 3-4x decode in about two weeks. There were numerous llama PRs for the feature too. Dozens of papers on arXiv to prove the concept. Crickets. Absolutely nothing. Well, except for a bunch of tl;dr: i'm going to close this because i'm too lazy to read it, lol. It's kind of impressive how dedicated to mediocrity llama.cpp maintainers are. So PSA: Use Strata, it's Qwen-3.8-flash-next on 8-16gb cards + 64gb ram, which is an almost Luna level model... and about as fast / faster than 27b (2000/70 t/s+)? Nice.

posted Sat, 03 Oct 2026 14:05:00 GMTseen 1 time
open on reddit ↗ 💬 128