3 papers
cs.AI2026
Token Predictors Are Not Planners: Building Physically Grounded Causal Reasoners
Zheng Lu, Mingqi Gao, Qinlei Xie +8
Current benchmarks for embodied vision-language planning often favor linguistic next-token prediction over physically grounded next-state reasoning. This rewards models that mimic…
cs.RO2025
LLaViDA: A Large Language Vision Driving Assistant for Explicit Reasoning and Enhanced Trajectory Planning
Yudong Liu, Spencer Hallyburton, Jiwoo Kim +8
Trajectory planning is a fundamental yet challenging component of autonomous driving. End-to-end planners frequently falter under adverse weather, unpredictable human behavior, or…
cs.CV2025
TrackOcc: Camera-based 4D Panoptic Occupancy Tracking
Zhuoguang Chen, Kenan Li, Xiuyu Yang +3
Comprehensive and consistent dynamic scene understanding from camera input is essential for advanced autonomous systems. Traditional camera-based perception tasks like 3D object tr…