4 papers · 1 filter
Think3D: Thinking with Space for Spatial Reasoning
Zaibin Zhang, Yuhan Wu, Lianjie Jia +10
While Vision-Language Models (VLMs) excel at 2D visual understanding, they remain constrained by 2D-centric paradigm that severely limits genuine 3D spatial reasoning. To bridge th…
AD-H: Language-guided Autonomous Driving with Hierarchical Agents
Zaibin Zhang, Talas Fu, Shiyu Tang +4
Language-guided autonomous driving requires bridging a large abstraction gap between high-level natural-language instructions and low-level vehicle control. End-to-end approaches t…
Assessment of Multimodal Large Language Models in Alignment with Human Values
Zhelun Shi, Zhipin Wang, Hongxing Fan +7
Large Language Models (LLMs) aim to serve as versatile assistants aligned with human values, as defined by the principles of being helpful, honest, and harmless (hhh). However, in…
From GPT-4 to Gemini and Beyond: Assessing the Landscape of MLLMs on Generalizability, Trustworthiness and Causality through Four Modalities
Chaochao Lu, Chen Qian, Guodong Zheng +33
Multi-modal Large Language Models (MLLMs) have shown impressive abilities in generating reasonable responses with respect to multi-modal contents. However, there is still a wide ga…