Publications (12)
Text-To-Speech Data Augmentation for Low Resource Speech Recognition
Rodolfo Zevallos
Nowadays, the main problem of deep learning techniques used in the development of automatic speech recognition (ASR) models is the lack of transcribed data. The goal of this resear…
How Attention Shapes Emotion: A Comparative Study of Attention Mechanisms for Speech Emotion Recognition
Marc Casals-Salvador, Federico Costa, Rodolfo Zevallos +1
Speech Emotion Recognition (SER) plays a key role in advancing human-computer interaction. Attention mechanisms have become the dominant approach for modeling emotional speech due…
Optimizing ASR for Catalan-Spanish Code-Switching: A Comparative Analysis of Methodologies
Carlos Mena, Pol Serra, Jacobo Romero +6
Code-switching (CS), the alternating use of two or more languages, challenges automatic speech recognition (ASR) due to scarce training data and linguistic similarities. The lack o…
Giving Voice to the Constitution: Low-Resource Text-to-Speech for Quechua and Spanish Using a Bilingual Legal Corpus
John E. Ortega, Rodolfo Zevallos, Fabricio Carraro
We present a unified pipeline for synthesizing high-quality Quechua and Spanish speech for the Peruvian Constitution using three state-of-the-art text-to-speech (TTS) architectures…
Evaluating Self-Supervised Speech Representations for Indigenous American Languages
Chih-Chen Chen, William Chen, Rodolfo Zevallos +1
The application of self-supervision to speech representation learning has garnered significant interest in recent years, due to its scalability to large amounts of unlabeled data.…
Huqariq: A Multilingual Speech Corpus of Native Languages of Peru for Speech Recognition
Rodolfo Zevallos, Luis Camacho, Nelsi Melgarejo
The Huqariq corpus is a multilingual collection of speech from native Peruvian languages. The transcribed corpus is intended for the research and development of speech technologies…