11 papers
SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models
Sirun Li, Minghao Liu, Ling Dai +4
Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textual cues, yet existing benchmarks rarely reveal how they arbitra…
MindEdit-Bench: Benchmarking Object-Level Counterfactual Spatial Reasoning in VLMs from In-the-Wild Photos
Leyuan Yu, Xiao Tang, Minghao Liu +6
Benchmarks for vision-language models (VLMs) mostly test observational spatial reasoning: models describe relations already visible in the input. Existing what-if tasks typically v…
OmniHD-Scenes: A Next-Generation Multimodal Dataset for Autonomous Driving
Lianqing Zheng, Long Yang, Qunshu Lin +10
The rapid advancement of deep learning has intensified the need for comprehensive data for use by autonomous driving algorithms. High-quality datasets are crucial for the developme…
PIN: A Knowledge-Intensive Dataset for Paired and Interleaved Multimodal Documents
Junjie Wang, Yuxiang Zhang, Minghao Liu +19
Recent advancements in large multimodal models (LMMs) have leveraged extensive multimodal datasets to enhance capabilities in complex knowledge-driven tasks. However, persistent ch…
MetaOcc: Spatio-Temporal Fusion of Surround-View 4D Radar and Camera for 3D Occupancy Prediction with Dual Training Strategies
Long Yang, Lianqing Zheng, Wenjin Ai +8
Robust 3D occupancy prediction is essential for autonomous driving, particularly under adverse weather conditions where traditional vision-only systems struggle. While the fusion o…
What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning
Yuchang Zhu, Huazhen Zhong, Qunshu Lin +6
With the remarkable generative capabilities of large language models (LLMs), using LLM-generated data to train downstream models has emerged as a promising approach to mitigate dat…