2 papers
cs.CV2024
LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences
Hongyan Zhi, Peihao Chen, Junyan Li +6
Research on 3D Vision-Language Models (3D-VLMs) is gaining increasing attention, which is crucial for developing embodied AI within 3D scenes, such as visual navigation and embodie…
cs.CV2023
Contrastive Vision-Language Alignment Makes Efficient Instruction Learner
Lizhao Liu, Xinyu Sun, Tianhang Xiang +3
We study the task of extending the large language model (LLM) into a vision-language instruction-following model. This task is crucial but challenging since the LLM is trained on t…