2 papers
cs.CV2026
Investigating Relational Reasoning in VLMs
Adhithya Laxman Ravi Shankar Geetha, Aulia Kharis Rakhmasari, Haleema Ramzan +1
Vision-Language Models (VLMs) achieve strong performance in visual reasoning tasks, but it remains unclear whether they understand visual relations, or simply employ shortcuts such…
cs.RO2026
Memory Over Maps: 3D Object Localization Without Reconstruction
Rui Zhou, Xander Yap, Jianwen Cao +3
Target localization is a prerequisite for embodied tasks such as navigation and manipulation. Conventional approaches rely on constructing explicit 3D scene representations to enab…