4 papers · 1 filter
Training-Free Multi-Step Inference for Target Speaker Extraction
Zhenghai You, Ying Shi, Lantian Li +1
Target speaker extraction (TSE) aims to recover a target speaker's speech from a mixture using a reference utterance as a cue. Most TSE systems adopt conditional auto-encoder archi…
MT-HuBERT: Self-Supervised Mix-Training for Few-Shot Keyword Spotting in Mixed Speech
Junming Yuan, Ying Shi, Dong Wang +2
Few-shot keyword spotting aims to detect previously unseen keywords with very limited labeled samples. A pre-training and adaptation paradigm is typically adopted for this task. Wh…
Few-Shot Keyword Spotting from Mixed Speech
Junming Yuan, Ying Shi, LanTian Li +2
Few-shot keyword spotting (KWS) aims to detect unknown keywords with limited training samples. A commonly used approach is the pre-training and fine-tuning framework. While effecti…
Serialized Output Training by Learned Dominance
Ying Shi, Lantian Li, Shi Yin +2
Serialized Output Training (SOT) has showcased state-of-the-art performance in multi-talker speech recognition by sequentially decoding the speech of individual speakers. To addres…