activity
20172022
most citedTransformer-Transducer: End-to-End Speech Recognition with Self-Attention

66 citations · 156 across the 21 of their papers we have counts for

collaborators
Showing cs.SDShow all

6 papers · 1 filter

cs.SD2022

Dynamic Speech Endpoint Detection with Regression Targets

Dawei Liang, Hang Su, Tarun Singh +5

Interactive voice assistants have been widely used as input interfaces in various scenarios, e.g. on smart homes devices, wearables and on AR devices. Detecting the end of a speech…

cs.SD20211 cited

Transferring Voice Knowledge for Acoustic Event Detection: An Empirical Study

Dawei Liang, Yangyang Shi, Yun Wang +6

Detection of common events and scenes from audio is useful for extracting and understanding human contexts in daily life. Prior studies have shown that leveraging knowledge from a…

cs.SD20211 cited

Flexi-Transducer: Optimizing Latency, Accuracy and Compute forMulti-Domain On-Device Scenarios

Jay Mahadeokar, Yangyang Shi, Yuan Shangguan +7

Often, the storage and computational constraints of embeddeddevices demand that a single on-device ASR model serve multiple use-cases / domains. In this paper, we propose aFlexible…

cs.SD2021

Dissecting User-Perceived Latency of On-Device E2E Speech Recognition

Yuan Shangguan, Rohit Prabhavalkar, Hang Su +8

As speech-enabled devices such as smartphones and smart speakers become increasingly ubiquitous, there is growing interest in building automatic speech recognition (ASR) systems th…

cs.SD2021

Memory-efficient Speech Recognition on Smart Devices

Ganesh Venkatesh, Alagappan Valliappan, Jay Mahadeokar +4

Recurrent transducer models have emerged as a promising solution for speech recognition on the current and next generation smart devices. The transducer models provide competitive…

cs.SD2020

Emformer: Efficient Memory Transformer Based Acoustic Model For Low Latency Streaming Speech Recognition

Yangyang Shi, Yongqiang Wang, Chunyang Wu +5

This paper proposes an efficient memory transformer Emformer for low latency streaming speech recognition. In Emformer, the long-range history context is distilled into an augmente…