activity
20202022
most citedWeak-Attention Suppression For Transformer Based Speech Recognition

5 citations · 13 across the 13 of their papers we have counts for

collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2022

Biased Self-supervised learning for ASR

Florian L. Kreyssig, Yangyang Shi, Jinxi Guo +3

Self-supervised learning via masked prediction pre-training (MPPT) has shown impressive performance on a range of speech-processing tasks. This paper proposes a method to bias self…

cs.CL2022

Streaming parallel transducer beam search with fast-slow cascaded encoders

Jay Mahadeokar, Yangyang Shi, Ke Li +5

Streaming ASR with strict latency constraints is required in many speech recognition applications. In order to achieve the required latency, streaming ASR models sacrifice accuracy…

cs.CL2021

Collaborative Training of Acoustic Encoders for Speech Recognition

Varun Nagaraja, Yangyang Shi, Ganesh Venkatesh +3

On-device speech recognition requires training models of different sizes for deploying on devices with various computational budgets. When building such different models, we can be…

cs.CL2021

Dynamic Encoder Transducer: A Flexible Solution For Trading Off Accuracy For Latency

Yangyang Shi, Varun Nagaraja, Chunyang Wu +9

We propose a dynamic encoder transducer (DET) for on-device speech recognition. One DET model scales to multiple devices with different computation capacities without retraining or…

cs.CL2021

Contextualized Streaming End-to-End Speech Recognition with Trie-Based Deep Biasing and Shallow Fusion

Duc Le, Mahaveer Jain, Gil Keren +9

How to leverage dynamic contextual information in end-to-end speech recognition has remained an active research area. Previous solutions to this problem were either designed for sp…

cs.CL2020

Streaming Attention-Based Models with Augmented Memory for End-to-End Speech Recognition

Ching-Feng Yeh, Yongqiang Wang, Yangyang Shi +4

Attention-based models have been gaining popularity recently for their strong performance demonstrated in fields such as machine translation and automatic speech recognition. One m…