4 papers
SpatioLM: Towards General Physical Spatial Intelligence in Vision-Language Models
Jing Wu, Jianhua Wu, Jiayi Guan +5
Vision-Language Models (VLMs) perform well on commonsense reasoning tasks but struggle with visual spatial reasoning. Most existing solutions introduce extra 3D prior inputs or ext…
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios
Jiaqi Fan, Jianhua Wu, Jincheng Gao +4
Multimodal large language models (MLLMs) have shown satisfactory effects in many autonomous driving tasks. In this paper, MLLMs are utilized to solve joint semantic scene understan…
Hallucination Elimination and Semantic Enhancement Framework for Vision-Language Models in Traffic Scenarios
Jiaqi Fan, Jianhua Wu, Hongqing Chu +2
Large vision-language models (LVLMs) have demonstrated remarkable capabilities in multimodal understanding and generation tasks. However, these models occasionally generate halluci…
Prospective Role of Foundation Models in Advancing Autonomous Vehicles
Jianhua Wu, Bingzhao Gao, Jincheng Gao +8
With the development of artificial intelligence and breakthroughs in deep learning, large-scale Foundation Models (FMs), such as GPT, Sora, etc., have achieved remarkable results i…