3 papers
cs.AI2026
Automated Trajectory Evaluation for Mobile Agents via Step-Level Consequence Reasoning and Aggregation
Pengshuai Yang, Zijing Gao, Xue Yu +3
Evaluating language-guided mobile agents has recently shifted from rule-based to model-based approaches to achieve scalable and automated assessments. However, existing holistic ev…
cs.CL2026
Test-Time Scaling for Small VLMs on Multilingual Visual MCQ
Spiros Baxevanakis, Peng-Jian Yang
Test-time scaling (TTS) reliably improves reasoning in large language models, but whether it transfers to small open vision-language models remains unclear. We examine this on EXAM…
cs.AI2026
SeerGuard: A Safety Framework for Mobile GUI Agents via World Model Prediction
Xue Yu, Bo Yuan, Kailin Zhao +3
Mobile graphical user interface (GUI) agents have demonstrated remarkable capabilities in automating complex tasks, yet they introduce critical safety risks because a single errone…