3 papers
eess.AS2026
Zero-Shot TTS With Enhanced Audio Prompts: Bsc Submission For The 2026 Wildspoof Challenge TTS Track
Jose Giraldo, Alex Peiró-Lilja, Rodolfo Zevallos +1
We evaluate two non-autoregressive architectures, StyleTTS2 and F5-TTS, to address the spontaneous nature of in-the-wild speech. Our models utilize flexible duration modeling to im…
cs.CL2025
Optimizing ASR for Catalan-Spanish Code-Switching: A Comparative Analysis of Methodologies
Carlos Mena, Pol Serra, Jacobo Romero +6
Code-switching (CS), the alternating use of two or more languages, challenges automatic speech recognition (ASR) due to scarce training data and linguistic similarities. The lack o…
eess.AS2024
Enhancing Crowdsourced Audio for Text-to-Speech Models
José Giraldo, Martà Llopart-Font, Alex Peiró-Lilja +3
High-quality audio data is a critical prerequisite for training robust text-to-speech models, which often limits the use of opportunistic or crowdsourced datasets. This paper prese…