1 paper
Ziqiang Wang, Li Gu, Zhixiang Chi +4
Mobile agent benchmarks increasingly rely on LLM-based judges to evaluate task completion, yet the reliability of these judges on mobile agent trajectories remains largely unexamin…