4 papers
SEE: Structure-aware Exploring & Exploiting for Long-horizon GUI Agent Trajectory Synthesis
Zhuohang Fan, Beichen Zhang, Yuanfa Li +4
Graphical User Interface (GUI) agents powered by vision-language models hold promise for automating real-world mobile tasks. However, progress is limited by the lack of high-covera…
Xiaomi-GUI-0 Technical Report
Wanxia Cao, Chengzhen Duan, Pei Fu +29
Graphical user interface (GUI) agents build on vision-language models to complete user tasks end-to-end in real applications through interface actions such as tapping, swiping, tex…
SEASON: Mitigating Temporal Hallucination in Video Large Language Models via Self-Diagnostic Contrastive Decoding
Chang-Hsun Wu, Kai-Po Chang, Yu-Yang Sheng +3
Video Large Language Models (VideoLLMs) have shown remarkable progress in video understanding. However, these models still struggle to effectively perceive and exploit rich tempora…
GUI-PRA: Process Reward Agent for GUI Tasks
Tao Xiong, Xavier Hu, Yurun Chen +6
Graphical User Interface (GUI) Agents powered by Multimodal Large Language Models (MLLMs) show significant potential for automating tasks. However, they often struggle with long-ho…