activity
20242026
most citedSpiking Tucker Fusion Transformer for Audio-Visual Zero-Shot Learning

3 citations · 3 across the 5 of their papers we have counts for

collaborators

6 papers

cs.CV2026

DROSR: Degradation-Disentangled Representation for Real-World Omnidirectional Image Super-Resolution

Hongyu An, Xinfeng Zhang, Xu Fan +3

With the growing demand for immersive visual experiences, high-quality omnidirectional images (ODIs) have become increasingly important. However, limitations in imaging devices and…

cs.LG2025

Adaptive Redundancy Regulation for Balanced Multimodal Information Refinement

Zhe Yang, Wenrui Li, Hongtao Chen +3

Multimodal learning aims to improve performance by leveraging data from multiple sources. During joint multimodal training, due to modality bias, the advantaged modality often domi…

cs.CV2025

PAGCNet: A Pose-Aware and Geometry Constrained Framework for Panoramic Depth Estimation

Kanglin Ning, Ruzhao Chen, Penghong Wang +3

Explicitly modeling room background depth as a geometric constraint has proven effective for panoramic depth estimation. However, reconstructing this background depth for regular e…

cs.CV2025

Spiking Variational Graph Representation Inference for Video Summarization

Wenrui Li, Wei Han, Liang-Jian Deng +2

With the rise of short video content, efficient video summarization techniques for extracting key information have become crucial. However, existing methods struggle to capture the…

cs.CV2024

Spatio-Temporal Distortion Aware Omnidirectional Video Super-Resolution

Hongyu An, Xinfeng Zhang, Shijie Zhao +2

Omnidirectional videos (ODVs) provide an immersive visual experience by capturing the 360° scene. With the rapid advancements in virtual/augmented reality, metaverse, and generativ…

cs.MM20243 cited

Spiking Tucker Fusion Transformer for Audio-Visual Zero-Shot Learning

Wenrui Li, Penghong Wang, Ruiqin Xiong +1

The spiking neural networks (SNNs) that efficiently encode temporal sequences have shown great potential in extracting audio-visual joint feature representations. However, coupling…