Lesson learned. Don't blindly trust repos and make sure everything is stable for a long running (multi weeks) benchmark.
I posted previously my swe-verified django 100 tasks benchmark comparing different local models and quantization. No new models for now, but a fix in my evaluation workflow that was unfortunately not stable during the weeks/months of me using it. I redid the evaluation on all runs and here are some noticeable changes: Flash Next is still king, but the benefit of xhigh vs medium reasoning effort is now properly showing. Same for 3.8 27B (however in everyday tasks I personally still prefer using medium)
scorecomments5 sightings
first seen 2026-10-05 11:34 UTClast seen 2026-10-05 19:36 UTCscore then 55score now 54gained -1sightings 5
open on reddit ↗
💬 9