5 papers
Dismantling Pathological Shortcuts: A Causal Framework for Faithful LVLM Decoding
Liu Yu, Can Chen, Ping Kuang +3
Large Vision-Language Models (LVLMs) exhibit sophisticated reasoning but remain susceptible to object hallucination. Deviating from the prevailing attention intensity assumption, w…
Beyond Detection: A Structure-Aware Framework for Scene Text Tracking
Chenmin Yu, Liu Yu, Daiqing Wu +3
Modern visual object trackers show impressive results on general targets, yet their performance drops substantially when dealing with scene text. Although currently underexplored,…
DRS-GUI: Dynamic Region Search for Training-Free GUI Grounding
Yichao Liu, Huawen Shen, Liu Yu +3
GUI agents powered by Multimodal Large Language Models (MLLMs) have demonstrated impressive capability in understanding and executing user instructions. However, accurately groundi…
StyleTextGen: Style-Conditioned Multilingual Scene Text Generation
Zeyu Chen, Fangmin Zhao, Yan Shu +3
Style-conditioned scene text generation faces unique challenges in extracting precise text styles from complex backgrounds and maintaining fine-grained style consistency across cha…
Causally-Grounded Dual-Path Attention Intervention for Object Hallucination Mitigation in LVLMs
Liu Yu, Zhonghao Chen, Ping Kuang +4
Object hallucination remains a critical challenge in Large Vision-Language Models (LVLMs), where models generate content inconsistent with visual inputs. Existing language-decoder…