2 papers
cs.CV2025
Guiding the Inner Eye: A Framework for Hierarchical and Flexible Visual Grounded Reasoning
Zhaoyang Wei, Wenchao Ding, Yanchao Hao +1
Models capable of "thinking with images" by dynamically grounding their reasoning in visual evidence represent a major leap in multimodal AI. However, replicating and advancing thi…
cs.CV2024
Reverse Region-to-Entity Annotation for Pixel-Level Visual Entity Linking
Zhengfei Xu, Sijia Zhao, Yanchao Hao +6
Visual Entity Linking (VEL) is a crucial task for achieving fine-grained visual understanding, matching objects within images (visual mentions) to entities in a knowledge base. Pre…