2 papers
cs.CV2024
Proximity QA: Unleashing the Power of Multi-Modal Large Language Models for Spatial Proximity Analysis
Jianing Li, Xi Nan, Ming Lu +2
Multi-modal large language models (MLLMs) have demonstrated remarkable vision-language capabilities, primarily due to the exceptional in-context understanding and multi-task learni…
cs.CV2023
SGL: Structure Guidance Learning for Camera Localization
Xudong Zhang, Shuang Gao, Xiaohu Nan +7
Camera localization is a classical computer vision task that serves various Artificial Intelligence and Robotics applications. With the rapid developments of Deep Neural Networks (…