3 papers
cs.CV2024
3DMIT: 3D Multi-modal Instruction Tuning for Scene Understanding
Zeju Li, Chao Zhang, Xiaoyan Wang +4
The remarkable potential of multi-modal large language models (MLLMs) in comprehending both vision and language information has been widely acknowledged. However, the scarcity of 3…
cs.CV2023
ClusVPR: Efficient Visual Place Recognition with Clustering-based Weighted Transformer
Yifan Xu, Pourya Shamsolmoali, Jie Yang
Visual place recognition (VPR) is a highly challenging task that has a wide range of applications, including robot navigation and self-driving vehicles. VPR is particularly difficu…
cs.CV2023
Exploring Multi-Modal Contextual Knowledge for Open-Vocabulary Object Detection
Yifan Xu, Mengdan Zhang, Xiaoshan Yang +1
In this paper, we for the first time explore helpful multi-modal contextual knowledge to understand novel categories for open-vocabulary object detection (OVD). The multi-modal con…