41 citations · 164 across the 32 of their papers we have counts for
Showing cs.SDShow all
2 papers · 1 filter
cs.SD2024
The Solution for Temporal Sound Localisation Task of ICCV 1st Perception Test Challenge 2023
Yurui Huang, Yang Yang, Shou Chen +3
In this paper, we propose a solution for improving the quality of temporal sound localization. We employ a multimodal fusion approach to combine visual and audio features. High-qua…
cs.SD2024
3S-TSE: Efficient Three-Stage Target Speaker Extraction for Real-Time and Low-Resource Applications
Shulin He, Jinjiang liu, Hao Li +3
Target speaker extraction (TSE) aims to isolate a specific voice from multiple mixed speakers relying on a registerd sample. Since voiceprint features usually vary greatly, current…