6 papers
BiRQ: Bi-Level Self-Labeling Random Quantization for Self-Supervised Speech Recognition
Liuyuan Jiang, Xiaodong Cui, Brian Kingsbury +2
Speech is a rich signal, and labeled audio-text pairs are costly, making self-supervised learning essential for scalable representation learning. A core challenge in speech SSL is…
Heterogeneous Self-Supervised Acoustic Pre-Training with Local Constraints
Xiaodong Cui, A F M Saif, Brian Kingsbury +1
Self-supervised pre-training using unlabeled data is widely used in automatic speech recognition. In this paper, we propose a new self-supervised pre-training approach to dealing w…
Objective Soups: Multilingual Multi-Task Modeling for Speech Processing
A F M Saif, Lisha Chen, Xiaodong Cui +3
Training a single model for multilingual, multi-task speech processing (MSP) is severely hampered by conflicting objectives between tasks like speech recognition and translation. W…
Efficient First-Order Optimization on the Pareto Set for Multi-Objective Learning under Preference Guidance
Lisha Chen, Quan Xiao, Ellen Hidemi Fukuda +3
Multi-objective learning under user-specified preference is common in real-world problems such as multi-lingual speech recognition under fairness. In this work, we frame such a pro…
Bilevel Joint Unsupervised and Supervised Training for Automatic Speech Recognition
Xiaodong Cui, A F M Saif, Songtao Lu +4
In this paper, we propose a bilevel joint unsupervised and supervised training (BL-JUST) framework for automatic speech recognition. Compared to the conventional pre-training and f…
FERERO: A Flexible Framework for Preference-Guided Multi-Objective Learning
Lisha Chen, AFM Saif, Yanning Shen +1
Finding specific preference-guided Pareto solutions that represent different trade-offs among multiple objectives is critical yet challenging in multi-objective problems. Existing…