← r/LocalLLaMA
▲
43
-2
18👁
r/LocalLLaMA · u/jacek2023 · 13d ago

ggml-cpu: tiled mul_mat for k-quants by jbooth · Pull Request #27851 · ggml-org/llama.cpp

faster CPU prompt processing: "TL;DR: 3-7x faster CPU mul\_mat using VNNI with IMO minimal complexity"

49 0 43 10/3 06:32 10/5 11:59 UTC
scorecomments18 sightings
first seen 2026-10-03 06:32 UTClast seen 2026-10-05 11:59 UTCscore then 45score now 43gained -2sightings 18
open on reddit ↗ 💬 20