10 papers
A Cocktail-Party Benchmark: Multi-Modal dataset and Comparative Evaluation Results
Thai-Binh Nguyen, Katerina Zmolikova, Pingchuan Ma +3
We introduce the task of Multi-Modal Context-Aware Recognition (MCoRec) in the ninth CHiME Challenge, which addresses the cocktail-party problem of overlapping conversations in a s…
Adapting Language Balance in Code-Switching Speech
Enes Yavuz Ugan, Ngoc-Quan Pham, Alexander Waibel
Despite achieving impressive results on standard benchmarks, large foundational models still struggle against code-switching test cases. When data scarcity cannot be used as the us…
Bayesian Low-Rank Factorization for Robust Model Adaptation
Enes Yavuz Ugan, Ngoc-Quan Pham, Alexander Waibel
Large speech foundation models achieve strong performance across many domains, but they often require adaptation to handle local needs such as code-switching, where speakers mix la…
End-to-End Evaluation for Low-Latency Simultaneous Speech Translation
Christian Huber, Tu Anh Dinh, Carlos Mullov +10
The challenge of low-latency speech translation has recently draw significant interest in the research community as shown by several publications and shared tasks. Therefore, it is…
Streaming Non-Autoregressive Model for Accent Conversion and Pronunciation Improvement
Tuan-Nam Nguyen, Ngoc-Quan Pham, Seymanur Akti +1
We propose a first streaming accent conversion (AC) model that transforms non-native speech into a native-like accent while preserving speaker identity, prosody and improving pronu…
Weight Factorization and Centralization for Continual Learning in Speech Recognition
Enes Yavuz Ugan, Ngoc-Quan Pham, Alexander Waibel
Modern neural network based speech recognition models are required to continually absorb new data without re-training the whole system, especially in downstream applications using…