MiMo V2.6 Pro almost matched GPT-6 Astra at figuring out what a user actually wants
Introducing InferBench: A benchmark testing how well frontier LLMs infer a user's priorities from their instructions.
We tested 12 LLMs across 20 scenarios with 2.8k conversations to see which models understand the user's goals.
Each conversation has a simulated user with a private profile of their priorities. The assistant decides on the best option or can first ask clarifying questions.
When a user's priorities were clearly stated up front, the models picked the best option 89% of the time. If some priorities were initially unstated then accuracy went down to 60%.
The open weight models did better than I expected:
GPT-6 Astra: 76%
MiMo V2.6 Pro: 75%
Gemini 3.1 Pro: 69%
Astra always asked clarifying questions when preferences weren't stated and never did when they were (extremely impressive). By comparison, Grok 4.6 asked to clarify in 18/64 conversations.
However LLMs are still just as confident when they're wrong. 105/288 wrong decisions were submitted with >90% confidence.
The full report is linked here: https://laugh.so/research/inferbench/
What surprised you the most?