most citedPatch-level Sounding Object Tracking for Audio-Visual Question Answering

2 citations · 4 across the 11 of their papers we have counts for

collaborators

11 papers

cs.CV2024

Linguistics-Vision Monotonic Consistent Network for Sign Language Production

Xu Wang, Shengeng Tang, Peipei Song +3

Sign Language Production (SLP) aims to generate sign videos corresponding to spoken language sentences, where the conversion of sign Glosses to Poses (G2P) is the key step. Due to…

cs.SD2024

Temporal-Frequency State Space Duality: An Efficient Paradigm for Speech Emotion Recognition

Jiaqi Zhao, Fei Wang, Kun Li +4

Speech Emotion Recognition (SER) plays a critical role in enhancing user experience within human-computer interaction. However, existing methods are overwhelmed by temporal domain…

cs.CV2024

Sign-IDD: Iconicity Disentangled Diffusion for Sign Language Production

Shengeng Tang, Jiayi He, Dan Guo +3

Sign Language Production (SLP) aims to generate semantically consistent sign videos from textual statements, where the conversion from textual glosses to sign poses (G2P) is a cruc…

cs.CV2024

Dense Audio-Visual Event Localization under Cross-Modal Consistency and Multi-Temporal Granularity Collaboration

Ziheng Zhou, Jinxing Zhou, Wei Qian +3

In the field of audio-visual learning, most research tasks focus exclusively on short videos. This paper focuses on the more practical Dense Audio-Visual Event Localization (DAVEL)…

cs.MM20242 cited

Patch-level Sounding Object Tracking for Audio-Visual Question Answering

Zhangbin Li, Jinxing Zhou, Jing Zhang +3

Answering questions related to audio-visual scenes, i.e., the AVQA task, is becoming increasingly popular. A critical challenge is accurately identifying and tracking sounding obje…

cs.CV2024

Towards Pixel-Level Prediction for Gaze Following: Benchmark and Approach

Feiyang Liu, Dan Guo, Jingyuan Xu +4

Following the gaze of other people and analyzing the target they are looking at can help us understand what they are thinking, and doing, and predict the actions that may follow. E…