12 citations · 12 across the 2 of their papers we have counts for
6 papers
Are you speaking my languages? On spoken language adherence in multimodal LLMs
Hyungwon Kim, Kandarp Joshi, Lillian Zhou +2
While Large Language Model (LLM) based Automatic Speech Recognition (ASR) enables seamless multilingual use, models often misidentify the output language, compromising transcriptio…
Data Quality Issues in Multilingual Speech Datasets: The Need for Sociolinguistic Awareness and Proactive Language Planning
Mingfei Lau, Qian Chen, Yeming Fang +3
Our quality audit for three widely used public multilingual speech datasets - Mozilla Common Voice 17.0, FLEURS, and Vox Populi - shows that in some languages, these datasets suffe…
Accented Speech Recognition: Benchmarking, Pre-training, and Diverse Data
Alëna Aksënova, Zhehuai Chen, Chung-Cheng Chiu +8
Building inclusive speech recognition systems is a crucial step towards developing technologies that speakers of all language varieties can use. Therefore, ASR systems must work fo…
Neural Simultaneous Speech Translation Using Alignment-Based Chunking
Patrick Wilken, Tamer Alkhouli, Evgeny Matusov +1
In simultaneous machine translation, the objective is to determine when to produce a partial translation given a continuous stream of source words, with a trade-off between latency…
Cumulative Adaptation for BLSTM Acoustic Models
Markus Kitza, Pavel Golik, Ralf Schlüter +1
This paper addresses the robust speech recognition problem as an adaptation task. Specifically, we investigate the cumulative application of adaptation methods. A bidirectional Lon…
A comprehensive study of batch construction strategies for recurrent neural networks in MXNet
Patrick Doetsch, Pavel Golik, Hermann Ney
In this work we compare different batch construction methods for mini-batch training of recurrent neural networks. While popular implementations like TensorFlow and MXNet suggest a…