8 papers
EvoGUI: An Evolution-Aware Benchmark for GUI State-Transition Understanding
Yaohan Yang, Minglei Shi, Borui Zhang +2
GUI agents must reason about how actions transform interface states, but end-to-end success rates entangle this ability with perception, grounding, planning, and recovery. We intro…
Syll: Open-Source Personal Automation with Cross-Surface Execution
Bo Zhang, Borui Zhang, Chenghao Jiang +5
Personal AI agents must increasingly operate across APIs, shells, web surfaces, and desktop GUIs, yet many systems remain tuned to a single interface and offer limited support for…
BAMI: Training-Free Bias Mitigation in GUI Grounding
Borui Zhang, Bo Zhang, Bo Wang +6
GUI grounding is a critical capability for enabling GUI agents to execute tasks such as clicking and dragging. However, in complex scenarios like the ScreenSpot-Pro benchmark, exis…
AdaZoom-GUI: Adaptive Zoom-based GUI Grounding with Instruction Refinement
Siqi Pei, Liang Tang, Tiaonan Duan +9
GUI grounding is a critical capability for vision-language models (VLMs) that enables automated interaction with graphical user interfaces by locating target elements from natural…
SFTok: Bridging the Performance Gap in Discrete Tokenizers
Qihang Rao, Borui Zhang, Wenzhao Zheng +2
Recent advances in multimodal models highlight the pivotal role of image tokenization in high-resolution image generation. By compressing images into compact latent representations…
SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder
Minglei Shi, Haolin Wang, Borui Zhang +11
Visual generation grounded in Visual Foundation Model (VFM) representations offers a highly promising unified pathway for integrating visual understanding, perception, and generati…