8 citations · 14 across the 15 of their papers we have counts for
9 papers · 1 filter
DriveZero: End-to-End Driving Beyond Human Demonstrations
Hao He, Chengcheng Hu, Zirun Su +17
Most end-to-end autonomous-driving systems learn by imitating human driving logs, leaving their learned behavior constrained by the quality and behavioral coverage of the recorded…
RoGe: Novel View Synthesis via End-to-End Implicit Reconstruction and Generation
Xiaolei Lang, Ze Kang, Zehao Huang +1
Novel view synthesis from sparse inputs requires both geometric grounding from the observed views and generative priors of unobserved regions, motivating recent hybrid methods that…
Object Concepts Emerge from Motion
Boshi Li, Xiaohui Wang, Xiaoyang Wu +3
Object-centric visual representations are important for physical-world perception, but existing visual pretraining methods often capture semantic categories without preserving the…
Geometry-Grounded Unified 3D Perception for Autonomous Driving
Longfei Xu, Xiaohui Wang, Zehao Huang +4
Camera-based autonomous driving perception requires a shared representation that preserves metric 3D structure across synchronized multi-camera streams. However, existing image-bas…
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes?
Han Li, Si Liu, Zehao Huang +6
Multimodal Large Language Models (MLLMs) have achieved remarkable success across diverse expert-level tasks, but they still struggle with fundamental abilities that humans naturall…
Generative Lane Topology Reasoning via Autoregressive Model with Geometry Prior
Jiahui Fu, Zehao Huang, Han Li +2
Lane topology reasoning aims to construct a lane graph from onboard sensor observations. Existing methods follow a detection and association paradigm that treats each lane instance…