14 papers · 1 filter
Large Depth Completion Model from Sparse Observations
Zhu Yu, Zhengyi Zhao, Runmin Zhang +7
This work presents the Large Depth Completion Model (LDCM), a simple, effective, and robust framework for single-view metric depth estimation with sparse observations. Without rely…
Embodied3DBench: Benchmarking Low-Level Embodied Spatial Intelligence of Vision Language Models
Jiyao Zhang, Mingxu Zhang, Yitong Peng +8
Are current Vision Language Models (VLMs) ready to comprehend and reason about complex embodied interactions in 3D environments? We introduce Embodied3DBench, a robot-centric bench…
Towards Consistent Video Geometry Estimation
Zhu Yu, Jingnan Gao, Runmin Zhang +9
This work presents ViGeo, a feed-forward foundation model for recovering spatially dense and temporally consistent geometry from video sequences. Built upon a plain transformer arc…
RC-GeoCP: Geometric Consensus for 4D Radar-Camera Collaborative Perception
Xiaokai Bai, Lianqing Zheng, Runwei Guan +3
Collaborative perception (CP) extends sensing range through feature sharing, but most systems remain LiDAR-centric. Camera and 4D radar sensing combines dense semantics with lower-…
Rethinking Unsupervised Cross-modal Flow Estimation: Learning from Decoupled Optimization and Consistency Constraint
Runmin Zhang, Jialiang Wang, Si-Yuan Cao +4
This work presents DCFlow, a novel unsupervised cross-modal flow estimation framework that integrates a decoupled optimization strategy and a cross-modal consistency constraint. Un…
EDFFDNet: Towards Accurate and Efficient Unsupervised Multi-Grid Image Registration
Haokai Zhu, Bo Qu, Si-Yuan Cao +4
Previous deep image registration methods that employ single homography, multi-grid homography, or thin-plate spline often struggle with real scenes containing depth disparities due…