1 citations · 1 across the 4 of their papers we have counts for
4 papers · 1 filter
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios
Jiaqi Fan, Jianhua Wu, Jincheng Gao +4
Multimodal large language models (MLLMs) have shown satisfactory effects in many autonomous driving tasks. In this paper, MLLMs are utilized to solve joint semantic scene understan…
Hallucination Elimination and Semantic Enhancement Framework for Vision-Language Models in Traffic Scenarios
Jiaqi Fan, Jianhua Wu, Hongqing Chu +2
Large vision-language models (LVLMs) have demonstrated remarkable capabilities in multimodal understanding and generation tasks. However, these models occasionally generate halluci…
RAC3: Retrieval-Augmented Corner Case Comprehension for Autonomous Driving with Vision-Language Models
Yujin Wang, Quanfeng Liu, Jiaqi Fan +5
Understanding and addressing corner cases is essential for ensuring the safety and reliability of autonomous driving systems. Vision-language models (VLMs) play a crucial role in e…
CAT: LoCalization and IdentificAtion Cascade Detection Transformer for Open-World Object Detection
Shuailei Ma, Yuefeng Wang, Jiaqi Fan +4
Open-world object detection (OWOD), as a more general and challenging goal, requires the model trained from data on known objects to detect both known and unknown objects and incre…