2 papers
cs.CV2024
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios
Jiaqi Fan, Jianhua Wu, Jincheng Gao +4
Multimodal large language models (MLLMs) have shown satisfactory effects in many autonomous driving tasks. In this paper, MLLMs are utilized to solve joint semantic scene understan…
cs.CV2024
Prospective Role of Foundation Models in Advancing Autonomous Vehicles
Jianhua Wu, Bingzhao Gao, Jincheng Gao +8
With the development of artificial intelligence and breakthroughs in deep learning, large-scale Foundation Models (FMs), such as GPT, Sora, etc., have achieved remarkable results i…