activity
20182025
most citedRelative Positional Encoding for Speech Recognition and Direct Translation

1 citations · 2 across the 16 of their papers we have counts for

collaborators
Showing cs.CLShow all

19 papers · 1 filter

cs.CL2025

Adapting Language Balance in Code-Switching Speech

Enes Yavuz Ugan, Ngoc-Quan Pham, Alexander Waibel

Despite achieving impressive results on standard benchmarks, large foundational models still struggle against code-switching test cases. When data scarcity cannot be used as the us…

cs.CL2025

Bayesian Low-Rank Factorization for Robust Model Adaptation

Enes Yavuz Ugan, Ngoc-Quan Pham, Alexander Waibel

Large speech foundation models achieve strong performance across many domains, but they often require adaptation to handle local needs such as code-switching, where speakers mix la…

cs.CL2025

A Cocktail-Party Benchmark: Multi-Modal dataset and Comparative Evaluation Results

Thai-Binh Nguyen, Katerina Zmolikova, Pingchuan Ma +3

We introduce the task of Multi-Modal Context-Aware Recognition (MCoRec) in the ninth CHiME Challenge, which addresses the cocktail-party problem of overlapping conversations in a s…

cs.CL2025

Streaming Non-Autoregressive Model for Accent Conversion and Pronunciation Improvement

Tuan-Nam Nguyen, Ngoc-Quan Pham, Seymanur Akti +1

We propose a first streaming accent conversion (AC) model that transforms non-native speech into a native-like accent while preserving speaker identity, prosody and improving pronu…

cs.CL2025

Weight Factorization and Centralization for Continual Learning in Speech Recognition

Enes Yavuz Ugan, Ngoc-Quan Pham, Alexander Waibel

Modern neural network based speech recognition models are required to continually absorb new data without re-training the whole system, especially in downstream applications using…

cs.CL2025

PIER: A Novel Metric for Evaluating What Matters in Code-Switching

Enes Yavuz Ugan, Ngoc-Quan Pham, Leonard Bärmann +1

Code-switching, the alternation of languages within a single discourse, presents a significant challenge for Automatic Speech Recognition. Despite the unique nature of the task, pe…