5 papers · 1 filter
PointVG-R: Internalizing Geometric Reasoning in MLLMs for Precise Pointing Localization via Visual Chain of Thought
Ling Li, Bowen Liu, Zinuo Zhan +5
Pointing-based visual grounding requires models to precisely locate target objects by deciphering complex spatial relationships between the visual scene and pointing gestures. Trad…
VistaRef: Boosting Visual Spatial Orientation Awareness for Pointing-to-Object Detection
Ling Li, Zhizhen Cai, Xinkun Wu +4
Grounding deictic gestures in natural images is fundamental to AR and human-robot collaboration, providing a basis for seamless spatial interaction. While Transformer-based visual…
ActWorld: From Explorable to Interactive World Model via Action-Aware Memory
Zhexiao Xiong, Yizhi Song, Hao Kang +11
Interactive world models aim to simulate environment dynamics under real-time user actions. However, their action vocabulary is largely confined to navigation: most actions corresp…
UGG: Unified Generative Grasping
Jiaxin Lu, Hao Kang, Haoxiang Li +4
Dexterous grasping aims to produce diverse grasping postures with a high grasping success rate. Regression-based methods that directly predict grasping parameters given the object…
Flexible Visual Recognition by Evidential Modeling of Confusion and Ignorance
Lei Fan, Bo Liu, Haoxiang Li +2
In real-world scenarios, typical visual recognition systems could fail under two major causes, i.e., the misclassification between known classes and the excusable misbehavior on un…