4 papers · 1 filter
Distill What RGB Can Recover: Privileged 3D Evidence for RGB-Only Vision-Language Models
Yanbin Hu, Jin Cui, Jun Ye +4
3D scene understanding requires reasoning about entity existence, spatial layout, and object relations, yet RGB images alone often provide insufficient 3D cues. Existing 3D-VLMs co…
VideoGPA: Distilling Geometry Priors for 3D-Consistent Video Generation
Hongyang Du, Junjie Ye, Xiaoyan Cong +7
While recent video diffusion models (VDMs) produce visually impressive results, they fundamentally struggle to maintain 3D structural consistency, often resulting in object deforma…
SMART: Advancing Scalable Map Priors for Driving Topology Reasoning
Junjie Ye, David Paz, Hengyuan Zhang +5
Topology reasoning is crucial for autonomous driving as it enables comprehensive understanding of connectivity and relationships between lanes and traffic elements. While recent ap…
A Language Agent for Autonomous Driving
Jiageng Mao, Junjie Ye, Yuxi Qian +2
Human-level driving is an ultimate goal of autonomous driving. Conventional approaches formulate autonomous driving as a perception-prediction-planning framework, yet their systems…