4 papers
WaterVideoQA: ASV-Centric Perception and Rule-Compliant Reasoning via Multi-Modal Agents
Runwei Guan, Shaofeng Liang, Ningwei Ouyang +9
While autonomous navigation has achieved remarkable success in passive perception (e.g., object detection and segmentation), it remains fundamentally constrained by a void in knowl…
ConLA: Contrastive Latent Action Learning from Human Videos for Robotic Manipulation
Weisheng Dai, Kai Lan, Jianyi Zhou +5
Vision-Language-Action (VLA) models achieve preliminary generalization through pretraining on large scale robot teleoperation datasets. However, acquiring datasets that comprehensi…
Large Foundation Models for Trajectory Prediction in Autonomous Driving: A Comprehensive Survey
Wei Dai, Shengen Wu, Wei Wu +7
Trajectory prediction serves as a critical functionality in autonomous driving, enabling the anticipation of future motion paths for traffic participants such as vehicles and pedes…
Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding
Runwei Guan, Ningwei Ouyang, Tianhao Xu +10
Automated waterway environment perception is crucial for enabling unmanned surface vessels (USVs) to understand their surroundings and make informed decisions. Most existing waterw…