184 citations · 474 across the 17 of their papers we have counts for
20 papers · 1 filter
Linguistically Informed Evaluation of Multilingual ASR for African Languages
Fei-Yueh Chen, Lateef Adeleke, C. M. Downey
Word Error Rate (WER) mischaracterizes ASR models' performance for African languages by combining phonological, tone, and other linguistic errors into a single lexical error. By co…
DM-Codec: Distilling Multimodal Representations for Speech Tokenization
Md Mubtasim Ahasan, Md Fahim, Tasnim Mohiuddin +6
Recent advancements in speech-language models have yielded significant improvements in speech tokenization and synthesis. However, effectively mapping the complex, multidimensional…
Speech-to-Speech Translation For A Real-world Unwritten Language
Peng-Jen Chen, Kevin Tran, Yilin Yang +13
We study speech-to-speech translation (S2ST) that translates speech from one language into another language and focuses on building systems to support languages without standard te…
Simple and Effective Unsupervised Speech Translation
Changhan Wang, Hirofumi Inaguma, Peng-Jen Chen +5
The amount of labeled data to train models for speech tasks is limited for most languages, however, the data scarcity is exacerbated for speech translation which requires labeled d…
Unified Speech-Text Pre-training for Speech Translation and Recognition
Yun Tang, Hongyu Gong, Ning Dong +8
We describe a method to jointly pre-train speech and text in an encoder-decoder modeling framework for speech translation and recognition. The proposed method incorporates four sel…
textless-lib: a Library for Textless Spoken Language Processing
Eugene Kharitonov, Jade Copet, Kushal Lakhotia +8
Textless spoken language processing research aims to extend the applicability of standard NLP toolset onto spoken language and languages with few or no textual resources. In this p…