4 papers
TRADE: Transducer-Augmented Decoder for Speech LLM
Yun Tang, Shanil Puri, Shinji Watanabe +1
Speech Large Language Models (Speech LLMs) lack a principled mechanism for streaming inference: their label-synchronous generation has no acoustic-frame alignment, making real-time…
Chunk Based Speech Pre-training with High Resolution Finite Scalar Quantization
Yun Tang, Cindy Tseng
Low latency speech human-machine communication is becoming increasingly necessary as speech technology advances quickly in the last decade. One of the primary factors behind the ad…
Peeking Into The Future For Contextual Biasing
Ramaneswaran Selvakumar, Cindy Tseng, Eesung Kim +2
While end-to-end (E2E) automatic speech recognition (ASR) models excel at general transcription, they struggle to recognize rare or unseen named entities (e.g., contact names, loca…
Enhanced Hybrid Transducer and Attention Encoder Decoder with Text Data
Yun Tang, Eesung Kim, Vijendra Raj Apsingekar
A joint speech and text optimization method is proposed for hybrid transducer and attention-based encoder decoder (TAED) modeling to leverage large amounts of text corpus and enhan…