5 papers · 1 filter
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes?
Han Li, Si Liu, Zehao Huang +6
Multimodal Large Language Models (MLLMs) have achieved remarkable success across diverse expert-level tasks, but they still struggle with fundamental abilities that humans naturall…
Generative Lane Topology Reasoning via Autoregressive Model with Geometry Prior
Jiahui Fu, Zehao Huang, Han Li +2
Lane topology reasoning aims to construct a lane graph from onboard sensor observations. Existing methods follow a detection and association paradigm that treats each lane instance…
Unified Modeling of Lane and Lane Topology for Driving Scene Reasoning
Han Li, Yulu Gao, Si Liu +3
Autonomous vehicles need to perceive not only physical elements in the driving scene, such as lane lines and traffic lights, but also logical elements like lane centerlines and the…
RATopo: Improving Lane Topology Reasoning via Redundancy Assignment
Han Li, Shaofei Huang, Longfei Xu +3
Lane topology reasoning plays a critical role in autonomous driving by modeling the connections among lanes and the topological relationships between lanes and traffic elements. Mo…
Enhancing 3D Lane Detection and Topology Reasoning with 2D Lane Priors
Han Li, Zehao Huang, Zitian Wang +3
3D lane detection and topology reasoning are essential tasks in autonomous driving scenarios, requiring not only detecting the accurate 3D coordinates on lane lines, but also reaso…