3 papers
cs.CV2026
Point What You Mean: Visually Grounded Instruction Policy
Hang Yu, Juntu Zhao, Yufeng Liu +9
Vision-Language-Action (VLA) models align vision and language with embodied control, but their object referring ability remains limited when relying solely on text prompt, especial…
cs.CV2026
VideoAfford: Grounding 3D Affordance from Human-Object-Interaction Videos via Multimodal Large Language Model
Hanqing Wang, Mingyu Liu, Xiaoyu Chen +9
3D affordance grounding aims to highlight the actionable regions on 3D objects, which is crucial for robotic manipulation. Previous research primarily focused on learning affordanc…
cs.CV2026
VAT: Vision Action Transformer by Unlocking Full Representation of ViT
Wenhao Li, Chengwei Ma, Weixin Mao
In robot learning, Vision Transformers (ViTs) are standard for visual perception, yet most methods discard valuable information by using only the final layer's features. We argue t…