activity
20242026
collaborators
Showing cs.SDShow all

5 papers · 1 filter

cs.SD2026

Enhancing BEST-RQ Pseudo-Label Quality through Online Refinement for Automatic Speech Recognition

Jingjing Xu, Zijian Yang, Mohammad Zeineldeen +3

BEST-RQ is a simple and effective self-supervised training method for speech representation learning that performs well on automatic speech recognition (ASR) tasks. It generates ps…

cs.SD2026

Sequence-Level Unsupervised Training in Speech Recognition: A Theoretical Study

Zijian Yang, Jörg Barkoczi, Ralf Schlüter +1

Unsupervised speech recognition is a task of training a speech recognition model with unpaired data. To determine when and how unsupervised speech recognition can succeed, and how…

cs.SD2025

A Comparative Analysis on ASR System Combination for Attention, CTC, Factored Hybrid, and Transducer Models

Noureldin Bayoumi, Robin Schmitt, Tina Raissi +3

Combination approaches for speech recognition (ASR) systems cover structured sentence-level or word-based merging techniques as well as combination of model scores during beam sear…

cs.SD2025

Label-Context-Dependent Internal Language Model Estimation for CTC

Zijian Yang, Minh-Nghia Phan, Ralf Schlüter +1

Although connectionist temporal classification (CTC) has the label context independence assumption, it can still implicitly learn a context-dependent internal language model (ILM)…

cs.SD2025

Right Label Context in End-to-End Training of Time-Synchronous ASR Models

Tina Raissi, Ralf Schlüter, Hermann Ney

Current time-synchronous sequence-to-sequence automatic speech recognition (ASR) models are trained by using sequence level cross-entropy that sums over all alignments. Due to the…