13 citations · 75 across the 46 of their papers we have counts for
36 papers · 1 filter
Training-Free Multi-Step Inference for Target Speaker Extraction
Zhenghai You, Ying Shi, Lantian Li +1
Target speaker extraction (TSE) aims to recover a target speaker's speech from a mixture using a reference utterance as a cue. Most TSE systems adopt conditional auto-encoder archi…
MT-HuBERT: Self-Supervised Mix-Training for Few-Shot Keyword Spotting in Mixed Speech
Junming Yuan, Ying Shi, Dong Wang +2
Few-shot keyword spotting aims to detect previously unseen keywords with very limited labeled samples. A pre-training and adaptation paradigm is typically adopted for this task. Wh…
An Investigation on Speaker Augmentation for End-to-End Speaker Extraction
Zhenghai You, Zhenyu Zhou, Lantian Li +1
Target confusion, defined as occasional switching to non-target speakers, poses a key challenge for end-to-end speaker extraction (E2E-SE) systems. We argue that this problem is la…
AlignVSR: Audio-Visual Cross-Modal Alignment for Visual Speech Recognition
Zehua Liu, Xiaolou Li, Chen Chen +3
Visual Speech Recognition (VSR) aims to recognize corresponding text by analyzing visual information from lip movements. Due to the high variability and weak information of lip mov…
Neural Scoring: A Refreshed End-to-End Approach for Speaker Recognition in Complex Conditions
Wan Lin, Junhui Chen, Tianhao Wang +3
Modern speaker verification systems primarily rely on speaker embeddings, followed by verification based on cosine similarity between the embedding vectors of the enrollment and te…
Quantitative Analysis of Audio-Visual Tasks: An Information-Theoretic Perspective
Chen Chen, Xiaolou Li, Zehua Liu +2
In the field of spoken language processing, audio-visual speech processing is receiving increasing research attention. Key components of this research include tasks such as lip rea…