3 citations · 7 across the 5 of their papers we have counts for
6 papers · 1 filter
Anatomy of Industrial Scale Multilingual ASR
Francis McCann Ramirez, Luka Chkhetiani, Andrew Ehrenberg +14
This paper describes AssemblyAI's industrial-scale automatic speech recognition (ASR) system, designed to meet the requirements of large-scale, multilingual ASR serving various app…
Two-pass Endpoint Detection for Speech Recognition
Anirudh Raju, Aparna Khare, Di He +9
Endpoint (EP) detection is a key component of far-field speech recognition systems that assist the user through voice commands. The endpoint detector has to trade-off between accur…
Separator-Transducer-Segmenter: Streaming Recognition and Segmentation of Multi-party Speech
Ilya Sklyar, Anna Piunova, Christian Osendorfer
Streaming recognition and segmentation of multi-party conversations with overlapping speech is crucial for the next generation of voice assistant applications. In this work we addr…
Streaming Multi-speaker ASR with RNN-T
Ilya Sklyar, Anna Piunova, Yulan Liu
Recent research shows end-to-end ASR systems can recognize overlapped speech from multiple speakers. However, all published works have assumed no latency constraints during inferen…
Improving RNN-T ASR Accuracy Using Context Audio
Andreas Schwarz, Ilya Sklyar, Simon Wiesler
We present a training scheme for streaming automatic speech recognition (ASR) based on recurrent neural network transducers (RNN-T) which allows the encoder network to learn to exp…
Subword Regularization: An Analysis of Scalability and Generalization for End-to-End Automatic Speech Recognition
Egor Lakomkin, Jahn Heymann, Ilya Sklyar +1
Subwords are the most widely used output units in end-to-end speech recognition. They combine the best of two worlds by modeling the majority of frequent words directly and at the…