7 papers
EvoGUI: An Evolution-Aware Benchmark for GUI State-Transition Understanding
Yaohan Yang, Minglei Shi, Borui Zhang +2
GUI agents must reason about how actions transform interface states, but end-to-end success rates entangle this ability with perception, grounding, planning, and recovery. We intro…
Attention at Rest Stays at Rest: Breaking Visual Inertia for Cognitive Hallucination Mitigation
Boyang Gong, Yu Zheng, Fanye Kong +2
Like a body at rest that stays at rest, we find that visual attention in multimodal large language models (MLLMs) exhibits pronounced inertia, remaining largely static once settled…
BWTA: Accurate and Efficient Binarized Transformer by Algorithm-Hardware Co-design
Yifu Ding, Xianglong Liu, Shenghao Jin +2
Ultra low-bit quantization brings substantial efficiency for Transformer-based models, but the accuracy degradation and limited GPU support hinder its wide usage. In this paper, we…
AdaZoom-GUI: Adaptive Zoom-based GUI Grounding with Instruction Refinement
Siqi Pei, Liang Tang, Tiaonan Duan +9
GUI grounding is a critical capability for vision-language models (VLMs) that enables automated interaction with graphical user interfaces by locating target elements from natural…
SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs
Jiahui Wang, Zuyan Liu, Yongming Rao +1
Multimodal Large Language Models (MLLMs) are commonly derived by extending pre-trained Large Language Models (LLMs) with visual capabilities. In this work, we investigate how MLLMs…
Ola: Pushing the Frontiers of Omni-Modal Language Model
Zuyan Liu, Yuhao Dong, Jiahui Wang +4
Recent advances in large language models, particularly following GPT-4o, have sparked increasing interest in developing omni-modal models capable of understanding more modalities.…