3 papers
eess.AS2023
A Deliberation-based Joint Acoustic and Text Decoder
Sepand Mavandadi, Tara N. Sainath, Ke Hu +1
We propose a new two-pass E2E speech recognition model that improves ASR performance by training on a combination of paired data and unpaired text data. Previously, the joint acous…
cs.CL2023
Massively Multilingual Shallow Fusion with Large Language Models
Ke Hu, Tara N. Sainath, Bo Li +7
While large language models (LLM) have made impressive progress in natural language processing, it remains unclear how to utilize them in improving automatic speech recognition (AS…
cs.CL2022
Improving Deliberation by Text-Only and Semi-Supervised Training
Ke Hu, Tara N. Sainath, Yanzhang He +4
Text-only and semi-supervised training based on audio-only data has gained popularity recently due to the wide availability of unlabeled text and speech data. In this work, we prop…