6 papers · 1 filter
PhoniTale: Phonologically Grounded Mnemonic Generation for Typologically Distant Language Pairs
Sana Kang, Myeongseok Gwon, Su Young Kwon +4
Vocabulary acquisition poses a significant challenge for second-language (L2) learners, especially when learning typologically distant languages such as English and Korean, where p…
OleSpeech-IV: A Large-Scale Multispeaker and Multilingual Conversational Speech Dataset with Diverse Topics
Wei Chu, Yuanzhe Dong, Ke Tan +7
OleSpeech-IV dataset is a large-scale multispeaker and multilingual conversational speech dataset with diverse topics. The audio content comes from publicly-available English podca…
CoLMbo: Speaker Language Model for Descriptive Profiling
Massa Baali, Shuo Han, Syed Abdul Hannan +5
Speaker recognition systems are often limited to classification tasks and struggle to generate detailed speaker characteristics or provide context-rich descriptions. These models p…
Lost in Transcription, Found in Distribution Shift: Demystifying Hallucination in Speech Foundation Models
Hanin Atwany, Abdul Waheed, Rita Singh +2
Speech foundation models trained at a massive scale, both in terms of model and data size, result in robust systems capable of performing multiple speech tasks, including automatic…
On the Robust Approximation of ASR Metrics
Abdul Waheed, Hanin Atwany, Rita Singh +1
Recent advances in speech foundation models are largely driven by scaling both model size and data, enabling them to perform a wide range of tasks, including speech recognition. Tr…
What Do Speech Foundation Models Not Learn About Speech?
Abdul Waheed, Hanin Atwany, Bhiksha Raj +1
Understanding how speech foundation models capture non-verbal cues is crucial for improving their interpretability and adaptability across diverse tasks. In our work, we analyze se…