← r/LocalLLaMA
▲
20
+11
5👁
r/LocalLLaMA · u/jacek2023 · 23h ago

ggml-cuda: assign four GDN state columns per warp by SongXiaoXi · Pull Request #30087 · ggml-org/llama.cpp

Another day, another Qwen 3.x speedup (prompt processing this time). Soon your Qwen will read your entire project before you can blink!

|test|master t/s|PR t/s|change|
|:-|:-|:-|:-|
|pp512|3075.24|3243.60|\+5.5%|
|pp4096|3059.87|3221.08|\+5.3%|
|tg128|47.62|47.67|\+0.1%|

21 0 20 10/8 11:29 10/9 01:36 UTC
scorecomments5 sightings
first seen 2026-10-08 11:29 UTClast seen 2026-10-09 01:36 UTCscore then 9score now 20gained +11sightings 5
open on reddit ↗ 💬 7