most citedSpiking Tucker Fusion Transformer for Audio-Visual Zero-Shot Learning

3 citations · 3 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV2025

Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot Learning

Wenrui Li, Penghong Wang, Xingtao Wang +3

Audio-visual zero-shot learning (ZSL) has been extensively researched for its capability to classify video data from unseen classes during training. Nevertheless, current methodolo…

cs.CV2024

Discrepancy-Aware Attention Network for Enhanced Audio-Visual Zero-Shot Learning

RunLin Yu, Yipu Gong, Wenrui Li +2

Audio-visual Zero-Shot Learning (ZSL) has attracted significant attention for its ability to identify unseen classes and perform well in video classification tasks. However, modal…

cs.CV2024

Digging into Intrinsic Contextual Information for High-fidelity 3D Point Cloud Completion

Jisheng Chu, Wenrui Li, Xingtao Wang +3

The common occurrence of occlusion-induced incompleteness in point clouds has made point cloud completion (PCC) a highly-concerned task in the field of geometric processing. Existi…

cs.AI2024

SHMamba: Structured Hyperbolic State Space Model for Audio-Visual Question Answering

Zhe Yang, Wenrui Li, Guanghui Cheng

The Audio-Visual Question Answering (AVQA) task holds significant potential for applications. Compared to traditional unimodal approaches, the multi-modal input of AVQA makes featu…

cs.MM20243 cited

Spiking Tucker Fusion Transformer for Audio-Visual Zero-Shot Learning

Wenrui Li, Penghong Wang, Ruiqin Xiong +1

The spiking neural networks (SNNs) that efficiently encode temporal sequences have shown great potential in extracting audio-visual joint feature representations. However, coupling…