3 papers
cs.AI2026
GSAR: Goal-State-Anchor Rewards for Mobile GUI Agents with Self-Evolving Data Synthesis
Long Zhang, Yuhan Chen, Chaoran Zhang +7
Vision-Language Models (VLMs) based GUI agents stand to benefit significantly from online reinforcement learning (RL). However, their training is bottlenecked by two fundamental is…
cs.AI2026
Xiaomi-GUI-0 Technical Report
Wanxia Cao, Chengzhen Duan, Pei Fu +29
Graphical user interface (GUI) agents build on vision-language models to complete user tasks end-to-end in real applications through interface actions such as tapping, swiping, tex…
cs.CL2025
Attention Basin: Why Contextual Position Matters in Large Language Models
Zihao Yi, Delong Zeng, Zhenqing Ling +6
The performance of Large Language Models (LLMs) is significantly sensitive to the contextual position of information in the input. To investigate the mechanism behind this position…