158 citations · 693 across the 78 of their papers we have counts for
21 papers · 1 filter
TTS-Guided Training for Accent Conversion Without Parallel Data
Yi Zhou, Zhizheng Wu, Mingyang Zhang +2
Accent Conversion (AC) seeks to change the accent of speech from one (source) to another (target) while preserving the speech content and speaker identity. However, many AC approac…
I4U System Description for NIST SRE'20 CTS Challenge
Kong Aik Lee, Tomi Kinnunen, Daniele Colibro +23
This manuscript describes the I4U submission to the 2020 NIST Speaker Recognition Evaluation (SRE'20) Conversational Telephone Speech (CTS) Challenge. The I4U's submission was resu…
Generate, Discriminate and Contrast: A Semi-Supervised Sentence Representation Learning Framework
Yiming Chen, Yan Zhang, Bin Wang +2
Most sentence embedding techniques heavily rely on expensive human-annotated sentence pairs as the supervised signals. Despite the use of large-scale unlabeled data, the performanc…
token2vec: A Joint Self-Supervised Pre-training Framework Using Unpaired Speech and Text
Xianghu Yue, Junyi Ao, Xiaoxue Gao +1
Self-supervised pre-training has been successful in both text and speech processing. Speech and text offer different but complementary information. The question is whether we are a…
Self-Supervised Training of Speaker Encoder with Multi-Modal Diverse Positive Pairs
Ruijie Tao, Kong Aik Lee, Rohan Kumar Das +2
We study a novel neural architecture and its training strategies of speaker encoder for speaker recognition without using any identity labels. The speaker encoder is trained to ext…
Explicit Intensity Control for Accented Text-to-speech
Rui Liu, Haolin Zuo, De Hu +2
Accented text-to-speech (TTS) synthesis seeks to generate speech with an accent (L2) as a variant of the standard version (L1). How to control the intensity of accent in the proces…