From the 1 of 6 linked papers with an AI index.
6 papers
SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them
Yang Zhou, Zixuan Huang, Sunzhu Li +10
The paper presents SpatialCLI, a framework that teaches vision-language models to use specialist visual tools for spatial reasoning and then internalize those capabilities, dramati…
Toward Deep Representation Learning for Event-Enhanced Visual Autonomous Perception: the eAP Dataset
Jinghang Li, Shichao Li, Qing Lian +3
Recent visual autonomous perception systems achieve remarkable performances with deep representation learning. However, they fail in scenarios with challenging illumination.While e…
MultiPark: Multimodal Parking Transformer with Next-Segment Prediction
Han Zheng, Zikang Zhou, Guli Zhang +6
Parking accurately and safely in highly constrained spaces remains a critical challenge. Unlike structured driving environments, parking requires executing complex maneuvers such a…
GoIRL: Graph-Oriented Inverse Reinforcement Learning for Multimodal Trajectory Prediction
Muleilan Pei, Shaoshuai Shi, Lu Zhang +2
Trajectory prediction for surrounding agents is a challenging task in autonomous driving due to its inherent uncertainty and underlying multimodality. Unlike prevailing data-driven…
Learning better representations for crowded pedestrians in offboard LiDAR-camera 3D tracking-by-detection
Shichao Li, Peiliang Li, Qing Lian +2
Perceiving pedestrians in highly crowded urban environments is a difficult long-tail problem for learning-based autonomous perception. Speeding up 3D ground truth generation for su…
SEPT: Standard-Definition Map Enhanced Scene Perception and Topology Reasoning for Autonomous Driving
Muleilan Pei, Jiayao Shan, Peiliang Li +4
Online scene perception and topology reasoning are critical for autonomous vehicles to understand their driving environments, particularly for mapless driving systems that endeavor…