stuntd 0.1.2: local heads for multi-field decisions, and why one weak field decides how often you skip the model
A week ago I posted stuntd here, a proxy that learns your LLM's typed decisions and answers the confident ones with a small local head (~20ms GPU, ~60ms CPU). Thanks for the feedback last time :)
0.1.2 is out, the main thing is decisions with several fields, like category + urgency + needs_human.
First idea was to answer each field locally when its head is sure and ask the model for the rest. Dropped it, you pay for the whole model call anyway and a half local half model answer is a pain to debug. So it's all or nothing now: local only when every field is sure, otherwise the model answers and every field becomes training data.
Didn't expect how much that costs. On the support demo the heads alone are sure on 99.9%, 92% and 76% of tickets, but all three at once only on 72.7%, so the weakest field decides.
It also retrains itself now. auto_retrain kicks in after N new captures, the new head sits in shadow next to the model, goes live when it agrees long enough and back to shadow if it starts losing. Anthropic Messages learns too, and there's serve --lazy.
Code: https://github.com/bladedevoff/stuntd
Try it: https://huggingface.co/spaces/pollix/stuntd
Anyone else doing multi-field outputs locally, is it one weak field for you too?