20 citations · 168 across the 65 of their papers we have counts for
27 papers · 1 filter
SA-WavLM: Speaker-Aware Self-Supervised Pre-training for Mixture Speech
Jingru Lin, Meng Ge, Junyi Ao +2
It was shown that pre-trained models with self-supervised learning (SSL) techniques are effective in various downstream speech tasks. However, most such models are trained on singl…
RefXVC: Cross-Lingual Voice Conversion with Enhanced Reference Leveraging
Mingyang Zhang, Yi Zhou, Yi Ren +3
This paper proposes RefXVC, a method for cross-lingual voice conversion (XVC) that leverages reference information to improve conversion performance. Previous XVC works generally t…
An Exploration of Length Generalization in Transformer-Based Speech Enhancement
Qiquan Zhang, Hongxu Zhu, Xinyuan Qian +2
The use of Transformer architectures has facilitated remarkable progress in speech enhancement. Training Transformers using substantially long speech utterances is often infeasible…
Target Speech Diarization with Multimodal Prompts
Yidi Jiang, Ruijie Tao, Zhengyang Chen +2
Traditional speaker diarization seeks to detect ``who spoke when'' according to speaker characteristics. Extending to target speech diarization, we detect ``when target event occur…
Autoregressive Diffusion Transformer for Text-to-Speech Synthesis
Zhijun Liu, Shuai Wang, Sho Inoue +2
Audio language models have recently emerged as a promising approach for various audio generation tasks, relying on audio tokenizers to encode waveforms into sequences of discrete s…
How Do Neural Spoofing Countermeasures Detect Partially Spoofed Audio?
Tianchi Liu, Lin Zhang, Rohan Kumar Das +3
Partially manipulating a sentence can greatly change its meaning. Recent work shows that countermeasures (CMs) trained on partially spoofed audio can effectively detect such spoofi…