2 papers
cs.RO2026
SpikeVLA: Vision-Language-Action Models with Spiking Neural Networks
Ruiqi Song, Dujun Nie, Siyu Teng +7
Vision-Language-Action (VLA) models have become a dominant paradigm for embodied intelligence. However, most existing approaches are built on large-scale transformers, resulting in…
cs.CV2025
InsightDrive: Insight Scene Representation for End-to-End Autonomous Driving
Ruiqi Song, Xianda Guo, Yanlun Peng +3
Conventional end-to-end autonomous driving methods often rely on explicit global scene representations, which typically consist of 3D object detection, online mapping, and motion p…