activity
20192021
most citedPhoneme-Based Contextualization for Cross-Lingual Speech Recognition in End-to-End Models

7 citations · 12 across the 4 of their papers we have counts for

collaborators

5 papers

eess.AS20212 cited

Learning Word-Level Confidence For Subword End-to-End ASR

David Qiu, Qiujia Li, Yanzhang He +9

We study the problem of word-level confidence estimation in subword-based end-to-end (E2E) models for automatic speech recognition (ASR). Although prior works have proposed trainin…

cs.CL2021

Transformer Based Deliberation for Two-Pass Speech Recognition

Ke Hu, Ruoming Pang, Tara N. Sainath +1

Interactive speech recognition systems must generate words quickly while also producing accurate results. Two-pass models excel at these requirements by employing a first-pass deco…

cs.CL20203 cited

A Streaming On-Device End-to-End Model Surpassing Server-Side Conventional Model Quality and Latency

Tara N. Sainath, Yanzhang He, Bo Li +26

Thus far, end-to-end (E2E) models have not been shown to outperform state-of-the-art conventional models with respect to both quality, i.e., word error rate (WER), and latency, i.e…

eess.AS2020

Deliberation Model Based Two-Pass End-to-End Speech Recognition

Ke Hu, Tara N. Sainath, Ruoming Pang +1

End-to-end (E2E) models have made rapid progress in automatic speech recognition (ASR) and perform competitively relative to conventional models. To further improve the quality, a…

cs.CL20197 cited

Phoneme-Based Contextualization for Cross-Lingual Speech Recognition in End-to-End Models

Ke Hu, Antoine Bruguier, Tara N. Sainath +2

Contextual automatic speech recognition, i.e., biasing recognition towards a given context (e.g. user's playlists, or contacts), is challenging in end-to-end (E2E) models. Such mod…