8 papers
SpikeVLA: Vision-Language-Action Models with Spiking Neural Networks
Ruiqi Song, Dujun Nie, Siyu Teng +7
Vision-Language-Action (VLA) models have become a dominant paradigm for embodied intelligence. However, most existing approaches are built on large-scale transformers, resulting in…
DriveSplat: Unified Neural Gaussian Reconstruction for Dynamic Driving Scenes
Cong Wang, Ruiqi Song, Wei Tian +3
Reconstructing large-scale dynamic driving scenes remains challenging due to the coexistence of static environments with extreme depth variation and diverse dynamic actors exhibiti…
Adjacent-view Transformers for Supervised Surround-view Depth Estimation
Xianda Guo, Wenjie Yuan, Yunpeng Zhang +5
Depth estimation has been widely studied and serves as the fundamental step of 3D perception for robotics and autonomous driving. Though significant progress has been made in monoc…
Stereo Anything: Unifying Zero-shot Stereo Matching with Large-Scale Mixed Data
Xianda Guo, Chenming Zhang, Youmin Zhang +8
Stereo matching serves as a cornerstone in 3D vision, aiming to establish pixel-wise correspondences between stereo image pairs for depth recovery. Despite remarkable progress driv…
StereoCarla: A High-Fidelity Driving Dataset for Generalizable Stereo
Xianda Guo, Chenming Zhang, Ruilin Wang +6
Stereo matching plays a crucial role in enabling depth perception for autonomous driving and robotics. While recent years have witnessed remarkable progress in stereo matching algo…
SURDS: Benchmarking Spatial Understanding and Reasoning in Driving Scenarios with Vision Language Models
Xianda Guo, Ruijun Zhang, Yiqun Duan +7
Accurate spatial reasoning in outdoor environments - covering geometry, object pose, and inter-object relationships - is fundamental to downstream tasks such as mapping, motion for…