3 papers
cs.CV2026
Beyond Language: Grounding Referring Expressions with Hand Pointing in Egocentric Vision
Ling Li, Bowen Liu, Zinuo Zhan +4
Traditional Visual Grounding (VG) predominantly relies on textual descriptions to localize objects, a paradigm that inherently struggles with linguistic ambiguity and often ignores…
cs.CV2026
How to Utilize Complementary Vision-Text Information for 2D Structure Understanding
Jiancheng Dong, Pengyue Jia, Derong Xu +9
LLMs typically linearize 2D tables into 1D sequences to fit their autoregressive architecture, which weakens row-column adjacency and other layout cues. In contrast, purely visual…
cs.CV2025
Mask What Matters: Controllable Text-Guided Masking for Self-Supervised Medical Image Analysis
Ruilang Wang, Shuotong Xu, Bowen Liu +3
The scarcity of annotated data in specialized domains such as medical imaging presents significant challenges to training robust vision models. While self-supervised masked image m…