6 papers · 1 filter
TUNI: Unifying Pre-training and Fine-tuning with Modality-Aware Mutual Learning and Rectification for RGB-T Semantic Segmentation
Xiaodong Guo, Xianda Guo, Tong Liu +4
RGB-thermal (RGB-T) semantic segmentation improves the environmental perception of autonomous platforms in challenging conditions. Prevailing RGB-T segmentation frameworks suffer f…
Adjacent-view Transformers for Supervised Surround-view Depth Estimation
Xianda Guo, Wenjie Yuan, Yunpeng Zhang +5
Depth estimation has been widely studied and serves as the fundamental step of 3D perception for robotics and autonomous driving. Though significant progress has been made in monoc…
Watch Where You Move: Region-aware Dynamic Aggregation and Excitation for Gait Recognition
Binyuan Huang, Yongdong Luo, Xianda Guo +4
Deep learning-based gait recognition has achieved great success in various applications. The key to accurate gait recognition lies in considering the unique and diverse behavior pa…
OpenStereo: A Comprehensive Benchmark for Stereo Matching and Strong Baseline
Xianda Guo, Chenming Zhang, Juntao Lu +5
Stereo matching aims to estimate the disparity between matching pixels in a stereo image pair, which is important to robotics, autonomous driving, and other computer vision tasks.…
MaskFuser: Masked Fusion of Joint Multi-Modal Tokenization for End-to-End Autonomous Driving
Yiqun Duan, Xianda Guo, Zheng Zhu +3
Current multi-modality driving frameworks normally fuse representation by utilizing attention between single-modality branches. However, the existing networks still suppress the dr…
Multi-Prompt with Depth Partitioned Cross-Modal Learning
Yingjie Tian, Yiqi Wang, Xianda Guo +2
In recent years, soft prompt learning methods have been proposed to fine-tune large-scale vision-language pre-trained models for various downstream tasks. These methods typically c…