Unsloth, Swift1.5, Peculiar-Ragdoll, ThinkingCap - Qwen3.8-27B
In a previous post I shared comparison between Swift1.5 and peculiar-ragdoll's checkpoints. Added the original unsloth Q4\_K\_XL and ThinkingCap Q4\_K\_M (they don't offer L or XL) to the comparison. Here are the results over a 69 set of eval questions.
All tests are now run at same "medium" reasoning effort.
unsloth-ud\_q4\_k\_xl one ran using llama.cpp - not the splash forked inference engine.
https://preview.redd.it/tztb66wygvsh1.png?width=2958&format=png&auto=…
I'll do a 3x repeat for the slow run to see if it maintains 69/69 each time.
EDIT: u/jucabala457 asked I test mradermacher/Signal-3.8-27B-Terse-Coder-i1-GGUF The Q4\_K\_M is closest quant available. A nice addition for sure! That GGUF couldn't run with Splash-based engine due to tensor incompat. I ran it using llama.cpp the slow way. The total time taken isn't a fair comparison for that reason. Updated results below
https://preview.redd.it/pvfvgd35twsh1.png?width=2976&format=png&auto=…
I also just made the tuieval tool available here https://github.com/ashe-wb/tuieval
Can't promise you the tool will work right away on your install since a fully local binary is what I've been using and testing with. Customize it with packs of domain-specific eval questions you deal with on the daily. This is the most important part. A model or fine-tune that is not good for one thing might be excellent for something else and only you know what your domain interests are. The ability of a model to render game graphics means nothing to me but it means everything to someone else.
https://preview.redd.it/m2n4rzlzjwsh1.png?width=2000&format=png&auto=…