collaborators

6 papers

cs.CL2025

BiRQ: Bi-Level Self-Labeling Random Quantization for Self-Supervised Speech Recognition

Liuyuan Jiang, Xiaodong Cui, Brian Kingsbury +2

Speech is a rich signal, and labeled audio-text pairs are costly, making self-supervised learning essential for scalable representation learning. A core challenge in speech SSL is…

cs.LG2025

Heterogeneous Self-Supervised Acoustic Pre-Training with Local Constraints

Xiaodong Cui, A F M Saif, Brian Kingsbury +1

Self-supervised pre-training using unlabeled data is widely used in automatic speech recognition. In this paper, we propose a new self-supervised pre-training approach to dealing w…

eess.AS2025

Objective Soups: Multilingual Multi-Task Modeling for Speech Processing

A F M Saif, Lisha Chen, Xiaodong Cui +3

Training a single model for multilingual, multi-task speech processing (MSP) is severely hampered by conflicting objectives between tasks like speech recognition and translation. W…

eess.AS2025

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities

George Saon, Avihu Dekel, Alexander Brooks +21

Granite-speech LLMs are compact and efficient speech language models specifically designed for English ASR and automatic speech translation (AST). The models were trained by modali…

cs.SD2025

A Non-autoregressive Model for Joint STT and TTS

Vishal Sunder, Brian Kingsbury, George Saon +5

In this paper, we take a step towards jointly modeling automatic speech recognition (STT) and speech synthesis (TTS) in a fully non-autoregressive way. We develop a novel multimoda…

cs.CL2024

Bilevel Joint Unsupervised and Supervised Training for Automatic Speech Recognition

Xiaodong Cui, A F M Saif, Songtao Lu +4

In this paper, we propose a bilevel joint unsupervised and supervised training (BL-JUST) framework for automatic speech recognition. Compared to the conventional pre-training and f…