← r/LocalLLaMA
▲
8
+1
10👁
r/LocalLLaMA · u/otacon6531 · 11d ago

IQ Quants still slow on P40?

I have been using Qwen 3.6:35b IQ4 via llama.cpp on my p40 and am getting anywhere between 37 - 83 tok/s (mtp is on). Prefill usually starts at 600 and slowly degrades as it continues processing so 600 for short prompts and more like 300-400 by the end of a long prompt. It hurts, but it is what my budget allows.

AI told me IQ quants are noticeably slower on the P40 and it referenced (https://www.reddit.com/r/LocalLLaMA/comments/1dmhpud/are\_iq\_quants\_slow\_o…) from two years ago, but I didn't feel it being slower when I moved from Q4 to IQ4, so...

What am I missing? Are IQ Quants actually a significant amount slower on the P40 or is this outdated information?

9 0 8 10/3 06:30 10/7 08:01 UTC
scorecomments10 sightings
first seen 2026-10-03 06:30 UTClast seen 2026-10-07 08:01 UTCscore then 7score now 8gained +1sightings 10
open on reddit ↗ 💬 5