1 citations · 1 across the 2 of their papers we have counts for
4 papers
Single-stage TTS with Masked Audio Token Modeling and Semantic Knowledge Distillation
Gerard I. Gállego, Roy Fejgin, Chunghsin Yeh +2
Audio token modeling has become a powerful framework for speech synthesis, with two-stage approaches employing semantic tokens remaining prevalent. In this paper, we aim to simplif…
Joint Semantic Knowledge Distillation and Masked Acoustic Modeling for Full-band Speech Restoration with Improved Intelligibility
Xiaoyu Liu, Xu Li, Joan Serrà +1
Speech restoration aims at restoring full-band speech with high quality and intelligibility, considering a diverse set of distortions. MaskSR is a recently proposed generative mode…
GASS: Generalizing Audio Source Separation with Large-scale Data
Jordi Pons, Xiaoyu Liu, Santiago Pascual +1
Universal source separation targets at separating the audio sources of an arbitrary mix, removing the constraint to operate on a specific domain like speech or music. Yet, the pote…
CLIPSonic: Text-to-Audio Synthesis with Unlabeled Videos and Pretrained Language-Vision Models
Hao-Wen Dong, Xiaoyu Liu, Jordi Pons +5
Recent work has studied text-to-audio synthesis using large amounts of paired text-audio data. However, audio recordings with high-quality text annotations can be difficult to acqu…