← r/LocalLLaMA
▲
9
 
8👁
r/LocalLLaMA · u/East-Muffin-6472 · 13d ago

My Reading Library: Evaluating LLMs on Android Tasks

post image

Can LLM agents actually get through a day in the life of a normal user? That question got me reading papers on Android agents and mobile benchmarks over the past few months. A few patterns kept showing up: - Most benchmarks run on emulators, making real-device metrics difficult to measure. - Important deployment metrics like battery, thermals, and temperature are often missing. - Everyday tasks are scattered across benchmarks, languages, and apps, rather than forming a consistent, globally relevant task set. - This makes it harder to evaluate whether an agent can actually work reliably on a real phone, for real users. For now, I’ve put together a library of papers on benchmarking mobile/Android agents for you all to read! Link: https://www.alphaxiv.org/shared/folder/01a070c6-29a0-77a9-a5b4-b670d5eee169

10 0 9 10/3 06:31 10/5 01:52 UTC
scorecomments8 sightings
first seen 2026-10-03 06:31 UTClast seen 2026-10-05 01:52 UTCscore then 9score now 9gained 0sightings 8
open on reddit ↗ 💬 2