4 papers
SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models
Sirun Li, Minghao Liu, Ling Dai +4
Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textual cues, yet existing benchmarks rarely reveal how they arbitra…
Geo-Embed: Towards Unified Multimodal Embeddings for Urban Understanding
Jiapeng Li, Yong Li, Junjie Zhou +2
Geospatial and urban applications increasingly require models to compare heterogeneous evidence across street-view imagery, remote-sensing observations, text descriptions, region p…
A Dynamic Bus Lane Strategy for Integrated Management of Human-Driven and Autonomous Vehicles
Haoran Li, Zhenzhou Yuan, Rui Yue +4
This study introduces a dynamic bus lane (DBL) strategy, referred to as the dynamic bus priority lane (DBPL) strategy, designed for mixed traffic environments featuring both manual…
Learning Street View Representations with Spatiotemporal Contrast
Yong Li, Yingjing Huang, Gengchen Mai +1
Street view imagery is extensively utilized in representation learning for urban visual environments, supporting various sustainable development tasks such as environmental percept…