← r/LocalLLaMA
▲
58
-1
28👁
r/LocalLLaMA · u/Embarrassed_Soup_279 · 32d ago

ExLlamaV3 is underrated

I moght get shit on for posting this but, I feel like i don't see this being talked enough and it feels like such a waste of a good piece of software. Exl3 is incredible, albeit only if you have NVIDIA cards I think?

Exl3 quants are higher quality for its size, much lower KLD metrics, faster, all compared to llama.cpp just from a few personal sets of tests I like to give my local models (these are not benchmarks). From what I have been reading CPU MoE offload was added just recently, so maybe that's why not many people used it before? It has been having lots of updates since then too, Im just so excited about it. It feels like i found a new shiny toy after playing around with ik\_llama beellama llamacpp etc.

I have been using tabbyAPI exl3 backend + qwen 3.8 27b sc 6bpw H6 and qwen 3.8 flash next 4bpw as my daily drivers and it's incredible what it can do. I hope this software gets known to more people too. I'm not affiliated with them or anything. i judt wanted to share it. It's just so cool, please give it a try!!

63 0 58 10/3 04:49 10/7 19:16 UTC
scorecomments28 sightings
first seen 2026-10-03 04:49 UTClast seen 2026-10-07 19:16 UTCscore then 59score now 58gained -1sightings 28
open on reddit ↗ 💬 82