1 paper
Ziqiang Wang, Ziqiang Wan, Li Gu +5
Mobile agent benchmarks increasingly rely on LLM-based judges to evaluate task completion, yet the reliability of these judges on mobile agent trajectories remains largely unexamin…