26 papers
RoadBench: Benchmarking MLLMs on Fine-Grained Spatial Understanding and Reasoning under Urban Road Scenarios
Jun Zhang, Xin Zhang, Jie Feng +5
Multimodal large language models (MLLMs) have demonstrated powerful capabilities in general spatial understanding and reasoning. However, their fine-grained spatial understanding a…
UrbanWell: Benchmarking Multimodal Large Language Models for Spatio-Temporal Urban Wellbeing Analytics
Yanxin Xi, Xiang Su, Jie Feng +3
Understanding urban wellbeing from multimodal data requires integrating heterogeneous spatial and temporal signals, posing significant challenges for current multimodal large langu…
SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes
Tianhui Liu, Jie Feng, Zhiheng Zheng +6
Humans can effortlessly perceive spatial layouts, form cognitive representations, reason about spatial relations, and translate such reasoning into actions in everyday 3D environme…
ARMove: Learning to Predict Human Mobility through Agentic Reasoning
Chuyue Wang, Jie Feng, Yuxi Wu +2
Human mobility prediction is a critical task but remains challenging due to its complexity and variability across populations and regions. Recently, large language models (LLMs) ha…
Enhancing Local Life Service Recommendation with Agentic Reasoning in Large Language Model
Shiteng Cao, Xiaochong Lan, Yuwei Du +4
Local life service recommendation is distinct from general recommendation scenarios due to its strong living need-driven nature. Fundamentally, accurately identifying a user's imme…
CityLens: Evaluating Large Vision-Language Models for Urban Socioeconomic Sensing
Tianhui Liu, Hetian Pang, Xin Zhang +5
Understanding urban socioeconomic conditions through visual data is a challenging yet essential task for sustainable urban development and policy planning. In this work, we introdu…