3 papers
cs.RO2026
InternVLA-A1: Unifying Understanding, Generation and Action for Robotic Manipulation
Junhao Cai, Zetao Cai, Jiafei Cao +39
Prevalent Vision-Language-Action (VLA) models are typically built upon Multimodal Large Language Models (MLLMs) and demonstrate exceptional proficiency in semantic understanding, b…
cs.RO2025
SimpleVSF: VLM-Scoring Fusion for Trajectory Prediction of End-to-End Autonomous Driving
Peiru Zheng, Yun Zhao, Zhan Gong +2
End-to-end autonomous driving has emerged as a promising paradigm for achieving robust and intelligent driving policies. However, existing end-to-end methods still face significant…
cs.CV2024
SimpleBEV: Improved LiDAR-Camera Fusion Architecture for 3D Object Detection
Yun Zhao, Zhan Gong, Peiru Zheng +2
More and more research works fuse the LiDAR and camera information to improve the 3D object detection of the autonomous driving system. Recently, a simple yet effective fusion fram…