3 papers
cs.SD2025
Finding My Voice: Generative Reconstruction of Disordered Speech for Automated Clinical Evaluation
Karen Rosero, Eunjung Yeo, David R. Mortensen +3
We present ChiReSSD, a speech reconstruction framework that preserves children speaker's identity while suppressing mispronunciations. Unlike prior approaches trained on healthy ad…
cs.CL2025
CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset
Brian Yan, Injy Hamed, Shuichiro Shimizu +24
We present CS-FLEURS, a new dataset for developing and evaluating code-switched speech recognition and translation systems beyond high-resourced languages. CS-FLEURS consists of 4…
eess.AS2023
w2v-SELD: A Sound Event Localization and Detection Framework for Self-Supervised Spatial Audio Pre-Training
Orlem Lima dos Santos, Karen Rosero, Roberto de Alencar Lotufo
Sound Event Detection and Localization (SELD) constitutes a complex task that depends on extensive multichannel audio recordings with annotated sound events and their respective lo…