← r/LocalLLaMA
▲
66
-3
29👁
r/LocalLLaMA · u/pmttyji · 28d ago

CUDA/HIP: Flash Attention tuning (gfx1201) by pwilkin · Pull Request #28102 · ggml-org/llama.cpp

Nice pp improvements for RDNA4(R9700) & 3.5(RX 9060 XT, 8060S). More good numbers on large context.

PR has detailed benchmarks.

u/ilintar 👍

73 0 66 10/3 04:49 10/8 21:23 UTC
scorecomments29 sightings
first seen 2026-10-03 04:49 UTClast seen 2026-10-08 21:23 UTCscore then 69score now 66gained -3sightings 29
open on reddit ↗ 💬 16