activity
20192021
most citedContrastive Semi-supervised Learning for ASR

2 citations · 6 across the 7 of their papers we have counts for

collaborators

8 papers

eess.AS20211 cited

Streaming Transformer Transducer Based Speech Recognition Using Non-Causal Convolution

Yangyang Shi, Chunyang Wu, Dilin Wang +9

This paper improves the streaming transformer transducer for speech recognition by using non-causal convolution. Many works apply the causal convolution to improve streaming transf…

cs.SD20211 cited

Transferring Voice Knowledge for Acoustic Event Detection: An Empirical Study

Dawei Liang, Yangyang Shi, Yun Wang +6

Detection of common events and scenes from audio is useful for extracting and understanding human contexts in daily life. Prior studies have shown that leveraging knowledge from a…

cs.SD20211 cited

Flexi-Transducer: Optimizing Latency, Accuracy and Compute forMulti-Domain On-Device Scenarios

Jay Mahadeokar, Yangyang Shi, Yuan Shangguan +7

Often, the storage and computational constraints of embeddeddevices demand that a single on-device ASR model serve multiple use-cases / domains. In this paper, we propose aFlexible…

cs.CL2021

Dynamic Encoder Transducer: A Flexible Solution For Trading Off Accuracy For Latency

Yangyang Shi, Varun Nagaraja, Chunyang Wu +9

We propose a dynamic encoder transducer (DET) for on-device speech recognition. One DET model scales to multiple devices with different computation capacities without retraining or…

cs.CL20212 cited

Contrastive Semi-supervised Learning for ASR

Alex Xiao, Christian Fuegen, Abdelrahman Mohamed

Pseudo-labeling is the most adopted method for pre-training automatic speech recognition (ASR) models. However, its performance suffers from the supervised teacher model's degradin…

cs.CL20201 cited

Transformer in action: a comparative study of transformer-based acoustic models for large scale speech recognition applications

Yongqiang Wang, Yangyang Shi, Frank Zhang +4

In this paper, we summarize the application of transformer and its streamable variant, Emformer based acoustic model for large scale speech recognition applications. We compare the…