← r/LocalLLaMA
▲
80
-4
22👁
r/LocalLLaMA · u/Danmoreng · 20d ago

I tested Qwen3.8 27B IQ3_XXS (10.18GiB) vs Bonsai Ternary PQ2 (6.42GiB)

post image

I did a small test of the new hyped quantisation of Qwen3.8 vs the biggest quant which fits into my limited 16GB VRAM with decent context. The results are interesting.

Of course, the smaller file gives worse results. However they are not that far off. Unfortunately, this comes at the expense of even more tokens beeing used by the Bonsai model and thus much longer generation times.

Visually I prefer the IQ3\_XXS results, but see for yourself.

The test is by no means scientific - just few UI generation tasks for direct comparison on the same hardware. Also, I ran llama.cpp with MTP while the Bonsai model doesn't seem to have MTP which makes it even slower.

[](https://github.com/Danmoreng/qwen3-8-27b-iq3-xxs-vs-bonsai/blob/main/RESULTS.…)

|Metric|Qwen IQ3|Bonsai PQ2|
|:-|:-|:-|
|Tasks completed|4/4|4/4|
|Fixed assertions|20/20|20/20|
|Agent wall time|8:00|24:09|
|Output tokens|27,197|84,176|
|Weighted decode|83.59 tok/s|64.91 tok/s|
|Speculative acceptance|65.22% MTP|39.76% modified N-gram|
|Compactions|0|0|
|Length stops|0|1|

Across the complete suite, Qwen finished 3.02× faster and used 3.10× fewer output tokens.

Results:

https://danmoreng.github.io/qwen3-8-27b-iq3-xxs-vs-bonsai/

Repo:

https://github.com/Danmoreng/qwen3-8-27b-iq3-xxs-vs-bonsai

88 0 80 10/3 04:47 10/8 19:12 UTC
scorecomments22 sightings
first seen 2026-10-03 04:47 UTClast seen 2026-10-08 19:12 UTCscore then 84score now 80gained -4sightings 22
open on reddit ↗ 💬 96