3 papers
cs.CL2023
LAST: Scalable Lattice-Based Speech Modelling in JAX
Ke Wu, Ehsan Variani, Tom Bagby +1
We introduce LAST, a LAttice-based Speech Transducer library in JAX. With an emphasis on flexibility, ease-of-use, and scalability, LAST implements differentiable weighted finite s…
eess.AS2023
JEIT: Joint End-to-End Model and Internal Language Model Training for Speech Recognition
Zhong Meng, Weiran Wang, Rohit Prabhavalkar +7
We propose JEIT, a joint end-to-end (E2E) model and internal language model (ILM) training method to inject large-scale unpaired text into ILM during E2E training which improves ra…
eess.AS2022
UserLibri: A Dataset for ASR Personalization Using Only Text
Theresa Breiner, Swaroop Ramaswamy, Ehsan Variani +6
Personalization of speech models on mobile devices (on-device personalization) is an active area of research, but more often than not, mobile devices have more text-only data than…