Referring to specifically Unsloth's UD\_Q4\_K\_XL quant because that's going to be a question, and is relevant regardless.
It's old now. It's not great at coding medium sized or even small-ish projects. I wouldn't hand my codebase to it by any means. It hallucinates, like any other model. It's not perfect, by any means.
I don't like the fine tunes; they're almost all coding focused, and lose general assistant capability to a strange degree.
What it is good at is general, broad agentic action.
Set the lights in the living room, and bash into this machine to get a movie going.
Send my grocery list to my phone/watch when I get to walmart.
Remind me to clock in at work each day because it's becoming a problem.
Tell my husband to come here because I'm under the car and I can't drop this thing that I finally got in just the right spot but need a third hand to get this wrench in the correct spot.
Tell my dad about the pets or any one of my projects because I'm showing off your memory to him.
This is all shit that it can do, consistently. Sometimes it'll thrash a bit, but it stays on task and is fast enough that little mistakes are a non-issue.
I recently got a couple of Tesla P100s to power the model, and it's made this model, that I already had going at a pretty good clip on limited hardware, to run at speeds that are genuinely conversational. \~120-140 t/s generation, \~1000PP at 0ctx, \~700 by 13k.
It's more than smart enough to know how to do these tasks and, with the right sampler settings, thinks ridiculously efficiently for them. Actually awesome.
Fucker got a minecraft mod pack installed and running on a machine it wasn't even running on through the prism launcher appimage. Don't worry, it only has ssh towards computers on my tailnet. It's probably fine.
It's bad at holding a persona. My TTS model is trained on Paul Bettany's MCU Jarvis. When the model emits the right phrasing, it feels like magic. It doesn't do that very often. Gemma4 is great at that.
I am actually so scared that if they do release a Qwen4 model in this class (30-40B parmeters, 2-5b active) that it's going to be a coding focused, overthinking, genuinely capable but not at all fast, mess. Gemma4 26b feels like it would be so close if it wasn't dumb as rocks. As it stands, I pay anthropic for that shit coding shit. Maybe when I can run Flash next at good speed (current \~30-40 t/s rn with strata, but at q3. It's ok.) I can kick claude to the curb.
I want a model like what 3.6 35b is, but smarter. Just as fast, but thinks of the little things. Memory updates. I ask it to add something to my calendar, but maybe it also sets up a dedicated notification for the specific time separately. That'd be nice. Not more coding focused. There are so many coding assistant bots. They're great. But the general AI assistant isn't solved yet for the normal person. Maybe it could have better general knowledge, but I think an n-gram table would solve that. It said the P100s used ROCm the other night. Silly bot.