From the 1 of 4 linked papers with an AI index.
4 papers
Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs
Zixuan Huang, Yang Zhou, Kaixuan Wang +7
Recent Vision-Language-Action (VLA) models for autonomous driving (AD) increasingly utilize chain-of-thought (CoT) supervision to enhance the reasoning capabilities of their Vision…
SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them
Yang Zhou, Zixuan Huang, Sunzhu Li +10
The paper presents SpatialCLI, a framework that teaches vision-language models to use specialist visual tools for spatial reasoning and then internalize those capabilities, dramati…
Toward Deep Representation Learning for Event-Enhanced Visual Autonomous Perception: the eAP Dataset
Jinghang Li, Shichao Li, Qing Lian +3
Recent visual autonomous perception systems achieve remarkable performances with deep representation learning. However, they fail in scenarios with challenging illumination.While e…
Learning better representations for crowded pedestrians in offboard LiDAR-camera 3D tracking-by-detection
Shichao Li, Peiliang Li, Qing Lian +2
Perceiving pedestrians in highly crowded urban environments is a difficult long-tail problem for learning-based autonomous perception. Speeding up 3D ground truth generation for su…