Let small local models write both the tests and the code. The tests rejected a known-correct solution 77% of the time
Classic setup, built by the book: one call writes a contract, one writes tests, one writes the code, a script runs the tests against the code, a repair step patches whatever fails. Four local GGUF models (qwen3-1.7b, qwen2.5-coder-1.5b, llama-3.2-1b, smollm2-360m), 6 coding tasks, T=0.3, 785 calls on a GTX 1660 SUPER 6GB.
Then the boring check nobody does: fed every generated test suite a known-correct reference solution. 129 of 168 rejected it. 77%.
The tests weren't lazy either. Average mutation score 0.965, they caught almost every mechanically broken version of the code. Of the suites with a perfect 1.0, 81% still failed the correct answer. Thorough, confident, testing the wrong spec. Wrong, see the edit at the bottom.
| model | correct code rejected by its own tests |
|---|---|
| llama-3.2-1b | 0.92-1.0 |
| qwen2.5-coder-1.5b | 0.71-0.78 |
| qwen3-1.7b | 0.47-0.65 |
| smollm2-360m | 1.0 |
Favorite case: count vowels. The code forgot uppercase. The tests checked "AaEeIiOoUu" and expected 5. Correct is 10, the buggy code returns 5. Tests and bug shared the same misunderstanding, the check said PASS, repair never ran. Only a hand-written test with "HELLO" caught it.
Repair: 67 attempts went to repair, and 50 of them were already-correct code the tests had rejected. Of 17 real bugs it fixed 1. Never broke working code, credit where due.
And the code itself wasn't the weak part. A single sample already solved 0.833 of tasks, and plain resampling got 0.958 at 8 samples (counted solved if any sample passes the reference tests, so it's pass@8 and still needs a judge in real life). At a similar token budget, 4 plain samples matched the pipeline without repair, on fewer tokens. These small models write correct code far more often than correct tests.
Caveats: 6 tasks, models from 360M to 1.7B, one temperature. Bigger models write better tests, how much better this doesn't say.
What I do since: the check that decides comes from the spec or from examples a human wrote. A model can propose tests, it doesn't get to be the judge.
Report (English version), harness and metrics, my repo: https://github.com/Deadatreides/LLM-MEASUREMENTS/blob/main/experiments/experi…
Anyone running local coding agents with self-written tests as the gate? Ever fed them a known-good answer?
(not a native speaker, an LLM helped with the English)
Edit: u/RobWattx was right about the mutation score, I checked the saved runs. The 129 suites that rejected the reference: 59 had wrong asserts, 39 had syntax errors, 29 had no test functions at all, 2 crashed. Broken suites fail every mutant too, so they get a perfect mutation score for free (126 of 129). Suites that accepted the reference: mean mutation score 0.888. So the honest numbers: 42% of the suites did not run at all, and of the suites that did run, 60% rejected the correct solution. The struck paragraph above was wrong.