← r/LocalLLaMA
▲
8
 
1👁
r/LocalLLaMA · u/EmilPi · 2h ago

Just another purely open-weight models benchmark

Live results, coding benchmarks included, agentic benchmarks included, domains and other classifications filters included. (Spoiler: DeepSeek-V4-Vision-Exp rules, but other models have their rule areas): https://beta.locallm.top

Evaluated by domain experts (my friends mostly; coding part is evaluated solely by me), classified by domains/languages/intents (can be filtered on the home page), agentic coding benchmark included. This is only public part of the data (whoever has private queries and evaluated models on them sees the results differently).

\*\*Some of the plethora of current limitations\*\*: evaluations/questions coverage for the domains/newer models is incomplete and imbalanced, only part of the questions classified, UI/UX under-developed, focus was on small models and lower quants.

\*Also\*: everything runs on local 2xRTX 3090 + 128GB RAM PC, everyone can create queries and evaluate models' answers, new runs (especially agentic) slow to appear (see PC specs).

No LLM-as-judge on purpose (and I sometimes regret it).

Benchmarks currently being extended & evaluated: 1) Agentic coding 2) one-shot coding 3) Agentic retrieval 4) Agentic story-writing.

New models being evaluated: 1) Qwen3.8-27B (Uncensored). New models to be evaluated soon 1) GLM-5.3-Flash-UD-IQ3XSS 2) Mellum2.1-12B-A2.5B

If this looks interesting to you:

CALL FOR HELP

I ask for help from those who also want to build a community benchmark, developers and domain experts alike! Many things are going to be implemented sooner or later, and the agentic benchmarks are just the beginning.

I call for help in these areas: 1) compute resources 2) evaluations 3) development - basically everything. But any proposal and feedback is valuable!

I will add more info and answer questions in the comments section.

P.S. No AI used for writing this post (even though English is actually not my native language /s).

posted Fri, 09 Oct 2026 18:43:54 GMTseen 1 time
open on reddit ↗ 💬 9