collaborators

7 papers

cs.AI2026

Xiaomi-GUI-0 Technical Report

Wanxia Cao, Chengzhen Duan, Pei Fu +29

Graphical user interface (GUI) agents build on vision-language models to complete user tasks end-to-end in real applications through interface actions such as tapping, swiping, tex…

cs.AI2026

GAIA: A Data Flywheel System for Training GUI Test-Time Scaling Critic Models

Shaokang Wang, Pei Fu, Ruoceng Zhang +7

While Large Vision-Language Models (LVLMs) have significantly advanced GUI agents' capabilities in parsing textual instructions, interpreting screen content, and executing tasks, a…

cs.CV2026

UniTranslator: A Unified Multi-modal Framework for End-to-end In-Image Machine Translation

Jiahao Lyu, Pei Fu, Zhenhang Li +6

In-Image Machine Translation (IIMT) aims to translate scene text in an image and render the translated text back into the original regions while preserving the overall visual appea…

cs.CV2026

Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning

Shaojie Zhang, Pei Fu, Ruoceng Zhang +8

Autonomous graphical user interface (GUI) agents rely on accurate GUI grounding, which maps language instructions to on-screen coordinates, to execute user commands. However, curre…

cs.CV2026

IMTBench: A Multi-Scenario Cross-Modal Collaborative Evaluation Benchmark for In-Image Machine Translation

Jiahao Lyu, Pei Fu, Zhenhang Li +7

End-to-end In-Image Machine Translation (IIMT) aims to convert text embedded within an image into a target language while preserving the original visual context, layout, and render…

cs.CV2025

BTL-UI: Blink-Think-Link Reasoning Model for GUI Agent

Shaojie Zhang, Ruoceng Zhang, Pei Fu +8

In the field of AI-driven human-GUI interaction automation, while rapid advances in multimodal large language models and reinforcement fine-tuning techniques have yielded remarkabl…