3 papers
cs.RO2026
SpikeVLA: Vision-Language-Action Models with Spiking Neural Networks
Ruiqi Song, Dujun Nie, Siyu Teng +7
Vision-Language-Action (VLA) models have become a dominant paradigm for embodied intelligence. However, most existing approaches are built on large-scale transformers, resulting in…
cs.CV2025
DriveSplat: Unified Neural Gaussian Reconstruction for Dynamic Driving Scenes
Cong Wang, Ruiqi Song, Wei Tian +3
Reconstructing large-scale dynamic driving scenes remains challenging due to the coexistence of static environments with extreme depth variation and diverse dynamic actors exhibiti…
cs.CV2025
InsightDrive: Insight Scene Representation for End-to-End Autonomous Driving
Ruiqi Song, Xianda Guo, Yanlun Peng +3
Conventional end-to-end autonomous driving methods often rely on explicit global scene representations, which typically consist of 3D object detection, online mapping, and motion p…