CUDA/HIP: Flash Attention tuning (gfx1201) by pwilkin · Pull Request #28102 · ggml-org/llama.cpp
Nice pp improvements for RDNA4(R9700) & 3.5(RX 9060 XT, 8060S). More good numbers on large context.
PR has detailed benchmarks.
scorecomments29 sightings
first seen 2026-10-03 04:49 UTClast seen 2026-10-08 21:23 UTCscore then 69score now 66gained -3sightings 29
open on reddit ↗
💬 16