Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
CoAdapt-GUI: Joint Workflow Context and Policy Adaptation for Unseen GUI Applications
Linqiang Guo, Li Gu, Zihuan Jiang +8
Mobile GUI agents remain brittle when deployed to applications absent from source training. We study novel-app generalization under a limited target interaction budget and without…
cs.AI2026
Benchmarking LLM Judges for Mobile Agent Evaluation
Ziqiang Wang, Ziqiang Wan, Li Gu +5
Mobile agent benchmarks increasingly rely on LLM-based judges to evaluate task completion, yet the reliability of these judges on mobile agent trajectories remains largely unexamin…
cs.AI2026
ETR: Entropy Trend Reward for Efficient Chain-of-Thought Reasoning
Xuan Xiong, Huan Liu, Li Gu +4
Chain-of-thought (CoT) reasoning improves large language model performance on complex tasks, but often produces excessively long and inefficient reasoning traces. Existing methods…