I tested 20+ ways to make a cheap coding model act like an expensive one. Here's what worked and what didn't
Short version from pre-registered experiments on real repo commits (Haiku as the cheap agent, Sonnet as the strong one). Every protocol was committed to git before its run, and later experiments used repos the designs had never seen. https://preview.redd.it/6rasddvfmpth1.png?width=1991&format=png&auto=… What worked: \- A stronger model that only speaks up when the agent repeats mistakes: +7 successes in 63, \~1.3x the cost (an always-on advisor got +8 but cost 3.5x). \- Running the agent's change and reporting facts ("if this line became \pass\, all tests would still pass") beats giving advice: 35/42 vs 32/42, formatting regressions 10 -> 0. \- Your preferences, captured in your own words, carried into every later task (15/15 vs 0/15). Just restating them in the prompt took compliance from 40% to 90%. What didn't: \- Memory of code knowledge, generic checklists, rules learned from git history, routing between models, and clarifying questions. \- For a strong model, none of it raised success (45/45 with or without). Cheapest per solved task: Haiku + "conscience" \~$1.22, Sonnet alone \~$1.41. Everything is public: paper, protocols, failures and the tool (source-available, non-commercial licence; works with Claude Code, Codex and OMP). Repo: https://github.com/abdullahbalabel/mihad Paper: https://github.com/abdullahbalabel/mihad/blob/main/paper/MIHAD\_Research\_Paper\_EN\_v2.7.md Happy to answer questions, and criticism of the method is very welcome.