1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.SD2023★ 1 cited
DINO-VITS: Data-Efficient Zero-Shot TTS with Self-Supervised Speaker Verification Loss for Noise Robustness
Vikentii Pankov, Valeria Pronina, Alexander Kuzmin +7
We address zero-shot TTS systems' noise-robustness problem by proposing a dual-objective training for the speaker encoder using self-supervised DINO loss. This approach enhances th…
cs.CL2023
Improving End-to-End Speech Processing by Efficient Text Data Utilization with Latent Synthesis
Jianqiao Lu, Wenyong Huang, Nianzu Zheng +3
Training a high performance end-to-end speech (E2E) processing model requires an enormous amount of labeled speech data, especially in the era of data-centric artificial intelligen…