6 papers
DROSR: Degradation-Disentangled Representation for Real-World Omnidirectional Image Super-Resolution
Hongyu An, Xinfeng Zhang, Xu Fan +3
With the growing demand for immersive visual experiences, high-quality omnidirectional images (ODIs) have become increasingly important. However, limitations in imaging devices and…
PAGCNet: A Pose-Aware and Geometry Constrained Framework for Panoramic Depth Estimation
Kanglin Ning, Ruzhao Chen, Penghong Wang +3
Explicitly modeling room background depth as a geometric constraint has proven effective for panoramic depth estimation. However, reconstructing this background depth for regular e…
Adaptive Redundancy Regulation for Balanced Multimodal Information Refinement
Zhe Yang, Wenrui Li, Hongtao Chen +3
Multimodal learning aims to improve performance by leveraging data from multiple sources. During joint multimodal training, due to modality bias, the advantaged modality often domi…
Spiking Variational Graph Representation Inference for Video Summarization
Wenrui Li, Wei Han, Liang-Jian Deng +2
With the rise of short video content, efficient video summarization techniques for extracting key information have become crucial. However, existing methods struggle to capture the…
Spatio-Temporal Distortion Aware Omnidirectional Video Super-Resolution
Hongyu An, Xinfeng Zhang, Shijie Zhao +2
Omnidirectional videos (ODVs) provide an immersive visual experience by capturing the 360° scene. With the rapid advancements in virtual/augmented reality, metaverse, and generati…
Spiking Tucker Fusion Transformer for Audio-Visual Zero-Shot Learning
Wenrui Li, Penghong Wang, Ruiqin Xiong +1
The spiking neural networks (SNNs) that efficiently encode temporal sequences have shown great potential in extracting audio-visual joint feature representations. However, coupling…