6 papers · 1 filter
Apple-: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence
Runmao Yao, Kairui Hu, Yukang Cao +11
Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physical law. Yet existing benchmarks largely evaluate physical plausi…
S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence
Yalun Dai, Hao Li, Shulin Tian +10
Real-world spatial intelligence requires reasoning over a continuous and evolving 3D world, yet existing VLMs and tool-augmented agents largely remain tied to static, stateless inf…
SpatialBench: Is Your Spatial Foundation Model an All-Round Player?
Haosong Peng, Hao Li, Jiaqi Chen +10
While spatial foundation models have demonstrated impressive performance on standard datasets, a critical question remains: are they truly all-round players capable of generalizing…
Rethinking VLM Representation for VLA Initialization
Weifeng Lin, Siyuan Huang, Hao Li +5
Vision-Language-Action (VLA) models widely adopt pretrained Vision-Language Models (VLMs) as policy backbones, yet it remains unclear what kind of pretrained VLM representation is…
Unified Multimodal Models as Auto-Encoders
Zhiyuan Yan, Kaiqing Lin, Zongjian Li +10
Image-to-text (I2T) understanding and text-to-image (T2I) generation are two fundamental, important yet traditionally isolated multimodal tasks. Despite their intrinsic connection,…
Exploring the Causality of End-to-End Autonomous Driving
Jiankun Li, Hao Li, Jiangjiang Liu +6
Deep learning-based models are widely deployed in autonomous driving areas, especially the increasingly noticed end-to-end solutions. However, the black-box property of these model…