Comparing Qwen3.8-27B fine-tunes and baselining vs. frontier
TL;DR:
- For my use case, mradermacher/Signal-3.8-27B-Terse-Coder.i1-Q4\_K\_M proved out. Ymmv based on your domain-specific tests.
- Time taken to solve problems compared to frontier models is massive; especially if you, like me, run a potato. My current feasible model's KPI over the full eval set is 28x slower. Time gap might be significantly more forgiving for folks with better hardware.
About a week ago, I shared prelim tests comparing Qwen3.8-27B fine-tunes. I've since ran a multi-day comprehensive 469 domain-specific eval against some of these fine-tunes.
I ran them all with llama.cpp and same thinking settings. I also ran the full eval set on Opus 5.5 and Astra as well as partially on Qwen3.8-Flash-Next via OpenRouter. Some providers disclose quantization and others do not. Would be nice if OpenRouter made it mandatory to do so.
Here's what I'm calling the PTA index. This will differ by model, eval, and hardware on hand. But can be part of a grounding KPI to measure one's progress by.
https://preview.redd.it/y27wkaoha5uh1.png?width=2966&format=png&auto=…
- Astra had the lowest token usage - although it appears the provider masks reasoning.
- No model got 100% accuracy. Opus 5.5 got close and topped the list at 99.6%.
- Of my local models, mradermacher/Signal-3.8-27B-Terse-Coder.i1-Q4\_K\_M had the lower token usage and time to completion of tasks while achieving higher accuracy.
- A Q4\_K\_M fine-tune performing better than other L or XL is a nice find. It overthought to cut off only once and passed more tests than others. I read somewhere that the Signal-Terse-Coder is a combo of AgentionAI's Signal-3.8-27B and Shockem's Terse-Coder LoRA. I don't understand the mechanics here but some sort of magic must be going on under the hood.
- I wouldn't read too much into TTFTs of models on OpenRouter. They cache and I can't do really much about it to influence it.
Median tokens and more interesting info below. The number of questions each model overthought on is represented in "Cut off" column (my setting at max 16,384 tokens per question): 1 violation by Signal-Terse-Coder, 4 by Dirk, and 4 by Unsloth.
https://preview.redd.it/iucw41rzc5uh1.png?width=2984&format=png&auto=…
19 minutes vs 9 hours is wild; counterpoint: as wild as 0 privacy for frontier is when compared with \~100 for local. Better hardware is the normalizer.
I suspect Astra is appearing to get fewest tokens for 456 questions out of 469 by hiding its reasoning. The same question answered by Opus 5.5 and Astra shows reasoning of Astra is perhaps masked by the provider. Nevertheless, the Astra response is succinct with no code commentary and a valid pass to what was asked.
Opus 5.5
https://preview.redd.it/87cicfq1f5uh1.png?width=2846&format=png&auto=…
Astra
https://preview.redd.it/aztvm4nkf5uh1.png?width=2858&format=png&auto=…
It's not in the list but gpt-oss-120b \*shat the bed big time\* on my pandas/numpy tasks. Only domain specific evals can uncover cases like that and warn you which models/fine-tunes to steer clear of for particular tasks.
I also tried AgentionAI/Qwen3.8-27B-AP-Q4\_K\_XL and UkisAI/Swift-1.5-Qwen3.8-27B-Q4\_K\_L but I cut the run around 115 questions for both. A clear pattern had developed by then and didn't see a need to let the run go on to completion.
https://preview.redd.it/zhiqa0bqt5uh1.png?width=2990&format=png&auto=…
Initally, I made the tool for myself. But have since decoupled the engine from tests/data to make it extensible to other users' choice. It's available here https://github.com/ashe-wb/tuieval