5 papers
Few-Shot Synthetic Accented Speech for ASR Fine-Tuning: What Helps and When?
Yurii Halychanskyi, Nimet Beyza Bozdag, Mark Hasegawa-Johnson +2
Synthetic accented speech is a promising way to improve automatic speech recognition (ASR) when real accented recordings are scarce. We ask what makes such data useful for ASR fine…
SPARCLE: SPeaker-aware Aligned Representations via Contrastive Language Embeddings
Priyam Mazumdar, Yurii Halychanskyi, Steven Guo +2
Recent advances in speech synthesis have shifted from phoneme representations to direct grapheme modeling. While phonemes address the one-to-many mapping between text and acoustics…
Accent Conversion: A Problem-Driven Survey of Sociolinguistic and Technical Constraints
Yurii Halychanskyi, Jianfeng Steven Guo, Volodymyr Kindratenko
Accent conversion has rapidly progressed alongside growing interest in improving global cross-cultural communication. This survey presents an overview of the evolution of accent co…
FAC-FACodec: Controllable Zero-Shot Foreign Accent Conversion with Factorized Speech Codec
Yurii Halychanskyi, Cameron Churchwell, Yutong Wen +1
Previous accent conversion (AC) methods, including foreign accent conversion (FAC), lack explicit control over the degree of modification. Because accent modification can alter the…
Latent Diffusion Bridges for Unsupervised Musical Audio Timbre Transfer
Michele Mancusi, Yurii Halychanskyi, Kin Wai Cheuk +8
Music timbre transfer is a challenging task that involves modifying the timbral characteristics of an audio signal while preserving its melodic structure. In this paper, we propose…