activity
20182022
most citedIncremental Layer-wise Self-Supervised Learning for Efficient Speech Domain Adaptation On Device

4 citations · 14 across the 11 of their papers we have counts for

collaborators
Showing eess.ASShow all

7 papers · 1 filter

eess.AS2022

Streaming End-to-End Multilingual Speech Recognition with Joint Language Identification

Chao Zhang, Bo Li, Tara Sainath +4

Language identification is critical for many downstream tasks in automatic speech recognition (ASR), and is beneficial to integrate into multilingual end-to-end ASR as an additiona…

eess.AS2021

Fast Contextual Adaptation with Neural Associative Memory for On-Device Personalized Speech Recognition

Tsendsuren Munkhdalai, Khe Chai Sim, Angad Chandorkar +4

Fast contextual adaptation has shown to be effective in improving Automatic Speech Recognition (ASR) of rare words and when combined with an on-device personalized training, it can…

eess.AS20202 cited

A Better and Faster End-to-End Model for Streaming ASR

Bo Li, Anmol Gulati, Jiahui Yu +12

End-to-end (E2E) models have shown to outperform state-of-the-art conventional models for streaming speech recognition [1] across many dimensions, including quality (as measured by…

eess.AS2020

Cascaded encoders for unifying streaming and non-streaming ASR

Arun Narayanan, Tara N. Sainath, Ruoming Pang +5

End-to-end (E2E) automatic speech recognition (ASR) models, by now, have shown competitive performance on several benchmarks. These models are structured to either operate in strea…

eess.AS2020

Confidence Estimation for Attention-based Sequence-to-sequence Models for Speech Recognition

Qiujia Li, David Qiu, Yu Zhang +5

For various speech-related tasks, confidence scores from a speech recogniser are a useful measure to assess the quality of transcriptions. In traditional hidden Markov model-based…

eess.AS20203 cited

Towards Fast and Accurate Streaming End-to-End ASR

Bo Li, Shuo-yiin Chang, Tara N. Sainath +4

End-to-end (E2E) models fold the acoustic, pronunciation and language models of a conventional speech recognition model into one neural network with a much smaller number of parame…