activity
20242026
collaborators

10 papers

cs.RO2026

SpikeVLA: Vision-Language-Action Models with Spiking Neural Networks

Ruiqi Song, Dujun Nie, Siyu Teng +7

Vision-Language-Action (VLA) models have become a dominant paradigm for embodied intelligence. However, most existing approaches are built on large-scale transformers, resulting in…

cs.CV2026

ROVR-Open-Dataset: A Large-Scale Depth Dataset for Autonomous Driving

Xianda Guo, Ruijun Zhang, Yiqun Duan +9

Depth estimation is a fundamental component of spatial perception for autonomous driving and other unmanned systems operating in open urban environments. Existing depth datasets su…

cs.CV2025

InsightDrive: Insight Scene Representation for End-to-End Autonomous Driving

Ruiqi Song, Xianda Guo, Yanlun Peng +3

Conventional end-to-end autonomous driving methods often rely on explicit global scene representations, which typically consist of 3D object detection, online mapping, and motion p…

cs.CV2025

Adjacent-view Transformers for Supervised Surround-view Depth Estimation

Xianda Guo, Wenjie Yuan, Yunpeng Zhang +5

Depth estimation has been widely studied and serves as the fundamental step of 3D perception for robotics and autonomous driving. Though significant progress has been made in monoc…

cs.CV2025

Stereo Anything: Unifying Zero-shot Stereo Matching with Large-Scale Mixed Data

Xianda Guo, Chenming Zhang, Youmin Zhang +8

Stereo matching serves as a cornerstone in 3D vision, aiming to establish pixel-wise correspondences between stereo image pairs for depth recovery. Despite remarkable progress driv…

cs.CV2025

StereoCarla: A High-Fidelity Driving Dataset for Generalizable Stereo

Xianda Guo, Chenming Zhang, Ruilin Wang +6

Stereo matching plays a crucial role in enabling depth perception for autonomous driving and robotics. While recent years have witnessed remarkable progress in stereo matching algo…