← r/LocalLLaMA
▲
54
-1
5👁
r/LocalLLaMA · u/WonderRico · 15d ago

Lesson learned. Don't blindly trust repos and make sure everything is stable for a long running (multi weeks) benchmark.

post image

I posted previously my swe-verified django 100 tasks benchmark comparing different local models and quantization. No new models for now, but a fix in my evaluation workflow that was unfortunately not stable during the weeks/months of me using it. I redid the evaluation on all runs and here are some noticeable changes: Flash Next is still king, but the benefit of xhigh vs medium reasoning effort is now properly showing. Same for 3.8 27B (however in everyday tasks I personally still prefer using medium)

55 0 54 10/5 11:34 10/5 19:36 UTC
scorecomments5 sightings
first seen 2026-10-05 11:34 UTClast seen 2026-10-05 19:36 UTCscore then 55score now 54gained -1sightings 5
open on reddit ↗ 💬 9