2 papers
cs.RO2026
SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models
Zonghe Liu, Shanyuan Jie, Xiaoquan Sun +4
Vision-Language-Action (VLA) models have shown strong potential for general robot manipulation, but most existing models rely on 2D visual-language backbones and lack fine-grained…
cs.CL2025
MERLIN: A Testbed for Multilingual Multimodal Entity Recognition and Linking
Sathyanarayanan Ramamoorthy, Vishwa Shah, Simran Khanuja +5
This paper introduces MERLIN, a novel testbed system for the task of Multilingual Multimodal Entity Linking. The created dataset includes BBC news article titles, paired with corre…