4 papers
PointVG-R: Internalizing Geometric Reasoning in MLLMs for Precise Pointing Localization via Visual Chain of Thought
Ling Li, Bowen Liu, Zinuo Zhan +5
Pointing-based visual grounding requires models to precisely locate target objects by deciphering complex spatial relationships between the visual scene and pointing gestures. Trad…
VistaRef: Boosting Visual Spatial Orientation Awareness for Pointing-to-Object Detection
Ling Li, Zhizhen Cai, Xinkun Wu +4
Grounding deictic gestures in natural images is fundamental to AR and human-robot collaboration, providing a basis for seamless spatial interaction. While Transformer-based visual…
Beyond Language: Grounding Referring Expressions with Hand Pointing in Egocentric Vision
Ling Li, Bowen Liu, Zinuo Zhan +4
Traditional Visual Grounding (VG) predominantly relies on textual descriptions to localize objects, a paradigm that inherently struggles with linguistic ambiguity and often ignores…
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation
Ziyu Zhu, Xilin Wang, Yixuan Li +9
Embodied scene understanding requires not only comprehending visual-spatial information that has been observed but also determining where to explore next in the 3D physical world.…