My Reading Library: Evaluating LLMs on Android Tasks
Can LLM agents actually get through a day in the life of a normal user? That question got me reading papers on Android agents and mobile benchmarks over the past few months. A few patterns kept showing up: - Most benchmarks run on emulators, making real-device metrics difficult to measure. - Important deployment metrics like battery, thermals, and temperature are often missing. - Everyday tasks are scattered across benchmarks, languages, and apps, rather than forming a consistent, globally relevant task set. - This makes it harder to evaluate whether an agent can actually work reliably on a real phone, for real users. For now, I’ve put together a library of papers on benchmarking mobile/Android agents for you all to read! Link: https://www.alphaxiv.org/shared/folder/01a070c6-29a0-77a9-a5b4-b670d5eee169