← r/LocalLLaMA
▲
22
 
1👁
r/LocalLLaMA · u/ZestRocket · 4h ago

Qwen3.8-27B UD-IQ4_XS Heretic + MTP on a 16 GB card with 55-68 tok/s (24gb and 12gb versions available too)

I wanted the uncensored Qwen3.8-27B (llmfan46's Heretic build, MTP head preserved) on my RTX 4080 with MTP on (because I've been testing TONS of configs, and realized that would be the one)

And none of the published quants were built for that: the good IQ4\_XS doesn't leave room for MTP, and the ones that fit are 3-bit.

I was amazed by that. as 16gb is VERY common, so decided to quantize it to fit, work well and loose as less as possible in quality:

I copied Unsloth's per-tensor UD recipe onto llmfan46's BF16, made 12 / 16 / 24 GB versions, and measured KLD against a Q8\_0 of the same weights for every quant I could find.

Quality (mean KLD vs Q8\_0, wikitext-2 / llama.cpp source code, lower is better)

|Quant|GiB|Prose|Code|
|:-|:-|:-|:-|
|UD-Q5\_K\_XL (mine, 24 GB)|19.44|0.0045|0.0037|
|mradermacher i1-IQ4\_XS|14.26|0.0202|0.0149|
|UD-IQ4\_XS (mine, 16 GB)|13.27|0.0268|0.0192|
|llmfan46 Q3\_K\_M|13.48|0.0639|0.0456|
|mradermacher i1-IQ3\_M|11.89|0.0649|0.0465|
|UD-IQ3\_XXS (mine, 12 GB)|10.18|0.0904|0.0587|
|mradermacher i1-Q2\_K|10.12|0.1551|0.1017|

Speed on the 4080 (llama.cpp b11457, 15K-token prompt, 32K ctx, MTP 2 drafts, mean of 3 fixed seeds): UD-IQ4\_XS does 55 tok/s on code and 50 on prose, against 29 / 29 without MTP. i1-IQ3\_M is faster (70 / 54) at 2.4× the KLD, I prefer quality here over speed, but your call.

Things I didn't expect:

  • 2 draft tokens beat 3. On the same seeds at 40K: 68 vs 56 tok/s on code, 55 vs 36 on prose. Fewer rejected drafts.
  • 48K vs 40K produced the exact same tokens (identical draft-acceptance counts per seed) and was 22% slower. That's pure VRAM spill into shared memory, nothing else.
  • The 16 GB edge is brutal. The same 40K config gave 68, 51 and 42 tok/s on code depending only on whether the desktop was holding 1.2, 1.4 or 1.9 GB of VRAM. Opening a Chrome window mid-generation took decode to 0.2 tok/s. If you're on Windows with a monitor on the same card, watch "shared GPU memory" in Task Manager.
  • MTP head precision barely matters. q6\_K head vs IQ3\_S head: 83% vs 83% acceptance on code, 58% vs 54% on prose.

Repo with all three files, the commands, and the scripts (recipe extraction, dry-run verification, KLD, seeded speed bench): https://huggingface.co/codavidgarcia/Qwen3.8-27B-Uncensored-Heretic-MTP-UD-GGUF

The 12 GB and 24 GB context numbers are computed from llama.cpp's reported buffers, not measured on those cards. If you run them, let me know what you get!

Credit to llmfan46 for the weights, Unsloth for the recipes and mradermacher for the imatrix

Enjoy!

edit: added the KLD chart since a few people asked about the numbers, (lower is better)

https://preview.redd.it/pxp048081juh1.png?width=3200&format=png&auto=…

posted Fri, 09 Oct 2026 22:54:37 GMTseen 1 time
open on reddit ↗ 💬 13