← r/LocalLLaMA
▲
38
+28
20👁
r/LocalLLaMA · u/Gold-Bat-3225 · 3d ago

MiMo V2.6 Pro almost matched GPT-6 Astra at figuring out what a user actually wants

post image

Introducing InferBench: A benchmark testing how well frontier LLMs infer a user's priorities from their instructions.

We tested 12 LLMs across 20 scenarios with 2.8k conversations to see which models understand the user's goals.

Each conversation has a simulated user with a private profile of their priorities. The assistant decides on the best option or can first ask clarifying questions.

When a user's priorities were clearly stated up front, the models picked the best option 89% of the time. If some priorities were initially unstated then accuracy went down to 60%.

The open weight models did better than I expected:

GPT-6 Astra: 76%

MiMo V2.6 Pro: 75%

Gemini 3.1 Pro: 69%

Astra always asked clarifying questions when preferences weren't stated and never did when they were (extremely impressive). By comparison, Grok 4.6 asked to clarify in 18/64 conversations.

However LLMs are still just as confident when they're wrong. 105/288 wrong decisions were submitted with >90% confidence.

The full report is linked here: https://laugh.so/research/inferbench/

What surprised you the most?

39 0 38 10/7 01:49 10/9 07:48 UTC
scorecomments20 sightings
first seen 2026-10-07 01:49 UTClast seen 2026-10-09 07:48 UTCscore then 10score now 38gained +28sightings 20
open on reddit ↗ 💬 10 (+6)