activity
20242026
collaborators

6 papers

cs.CV2026

DROSR: Degradation-Disentangled Representation for Real-World Omnidirectional Image Super-Resolution

Hongyu An, Xinfeng Zhang, Xu Fan +3

With the growing demand for immersive visual experiences, high-quality omnidirectional images (ODIs) have become increasingly important. However, limitations in imaging devices and…

cs.CV2026

PAGCNet: A Pose-Aware and Geometry Constrained Framework for Panoramic Depth Estimation

Kanglin Ning, Ruzhao Chen, Penghong Wang +3

Explicitly modeling room background depth as a geometric constraint has proven effective for panoramic depth estimation. However, reconstructing this background depth for regular e…

cs.LG2025

Adaptive Redundancy Regulation for Balanced Multimodal Information Refinement

Zhe Yang, Wenrui Li, Hongtao Chen +3

Multimodal learning aims to improve performance by leveraging data from multiple sources. During joint multimodal training, due to modality bias, the advantaged modality often domi…

cs.CV2025

Spiking Variational Graph Representation Inference for Video Summarization

Wenrui Li, Wei Han, Liang-Jian Deng +2

With the rise of short video content, efficient video summarization techniques for extracting key information have become crucial. However, existing methods struggle to capture the…

cs.CV2025

Spatio-Temporal Distortion Aware Omnidirectional Video Super-Resolution

Hongyu An, Xinfeng Zhang, Shijie Zhao +2

Omnidirectional videos (ODVs) provide an immersive visual experience by capturing the 360° scene. With the rapid advancements in virtual/augmented reality, metaverse, and generati…

cs.MM2024

Spiking Tucker Fusion Transformer for Audio-Visual Zero-Shot Learning

Wenrui Li, Penghong Wang, Ruiqin Xiong +1

The spiking neural networks (SNNs) that efficiently encode temporal sequences have shown great potential in extracting audio-visual joint feature representations. However, coupling…