← r/LocalLLaMA
▲
12
+1
34👁
r/LocalLLaMA · u/klieret · 7d ago

New benchmark on LMs fixing bugs before users run into them

Hi! This is a new benchmark that I created together with other researchers at Meta, Stanford, Harvard, UW.

Most benchmarks these days seem to test models to just fix a bug that I as a user already encountered. But shouldn't we expect models by now to also find bugs before anyone runs into them?

So in SWE-sweep we just hand an agent a big codebase and ask it to find & fix as many bugs as it can. We then give a score based on a hidden set of bugs that we know about in the repos. All the bugs are real-world bugs. We do a lot of filtering to make sure the bugs are actually discoverable & fixable from reading the repo alone.

https://preview.redd.it/irffy7x5v2th1.png?width=1080&format=png&auto=…

We're still expanding the leaderboard list with more local models (unfortunately it's always a big harder with funding/infra etc), but right now it seems like it's quite hard to beat Luna xhigh in terms of cost efficiency.

Also the scores are way lower than I would've expected. Some tasks are legitimately superhuman in practice (like fixing up all of numpy), but there's also lots of small repos, where I would've expected a lot more from current models.

Everything is open source (MIT license) on github and you can find paper etc. on the website.

Happy to answer questions here, also super curious what open weights models you'd recommend running next (we're working on an update next week).

14 0 12 10/3 06:28 10/9 04:18 UTC
scorecomments34 sightings
first seen 2026-10-03 06:28 UTClast seen 2026-10-09 04:18 UTCscore then 11score now 12gained +1sightings 34
open on reddit ↗ 💬 14 (+2)