7 papers
Xiaomi-GUI-0 Technical Report
Wanxia Cao, Chengzhen Duan, Pei Fu +29
Graphical user interface (GUI) agents build on vision-language models to complete user tasks end-to-end in real applications through interface actions such as tapping, swiping, tex…
GAIA: A Data Flywheel System for Training GUI Test-Time Scaling Critic Models
Shaokang Wang, Pei Fu, Ruoceng Zhang +7
While Large Vision-Language Models (LVLMs) have significantly advanced GUI agents' capabilities in parsing textual instructions, interpreting screen content, and executing tasks, a…
UniTranslator: A Unified Multi-modal Framework for End-to-end In-Image Machine Translation
Jiahao Lyu, Pei Fu, Zhenhang Li +6
In-Image Machine Translation (IIMT) aims to translate scene text in an image and render the translated text back into the original regions while preserving the overall visual appea…
Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning
Shaojie Zhang, Pei Fu, Ruoceng Zhang +8
Autonomous graphical user interface (GUI) agents rely on accurate GUI grounding, which maps language instructions to on-screen coordinates, to execute user commands. However, curre…
IMTBench: A Multi-Scenario Cross-Modal Collaborative Evaluation Benchmark for In-Image Machine Translation
Jiahao Lyu, Pei Fu, Zhenhang Li +7
End-to-end In-Image Machine Translation (IIMT) aims to convert text embedded within an image into a target language while preserving the original visual context, layout, and render…
BTL-UI: Blink-Think-Link Reasoning Model for GUI Agent
Shaojie Zhang, Ruoceng Zhang, Pei Fu +8
In the field of AI-driven human-GUI interaction automation, while rapid advances in multimodal large language models and reinforcement fine-tuning techniques have yielded remarkabl…