3 papers
cs.CV2026
PAGCNet: A Pose-Aware and Geometry Constrained Framework for Panoramic Depth Estimation
Kanglin Ning, Ruzhao Chen, Penghong Wang +3
Explicitly modeling room background depth as a geometric constraint has proven effective for panoramic depth estimation. However, reconstructing this background depth for regular e…
cs.LG2025
Adaptive Redundancy Regulation for Balanced Multimodal Information Refinement
Zhe Yang, Wenrui Li, Hongtao Chen +3
Multimodal learning aims to improve performance by leveraging data from multiple sources. During joint multimodal training, due to modality bias, the advantaged modality often domi…
cs.CV2025
Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning
Wenrui Li, Penghong Wang, Xingtao Wang +3
Audio-visual zero-shot learning (ZSL) has been extensively researched for its capability to classify video data from unseen classes during training. Nevertheless, current methodolo…