6 papers
LT-LM: a novel non-autoregressive language model for single-shot lattice rescoring
Anton Mitrofanov, Mariya Korenevskaya, Ivan Podluzhny +7
Neural network-based language models are commonly used in rescoring approaches to improve the quality of modern automatic speech recognition (ASR) systems. Most of the existing met…
Exploration of End-to-End ASR for OpenSTT -- Russian Open Speech-to-Text Dataset
Andrei Andrusenko, Aleksandr Laptev, Ivan Medennikov
This paper presents an exploration of end-to-end automatic speech recognition systems (ASR) for the largest open-source Russian language data set -- OpenSTT. We evaluate different…
Target-Speaker Voice Activity Detection: a Novel Approach for Multi-Speaker Diarization in a Dinner Party Scenario
Ivan Medennikov, Maxim Korenevsky, Tatiana Prisyach +9
Speaker diarization for real-life scenarios is an extremely challenging problem. Widely used clustering-based diarization approaches perform rather poorly in such conditions, mainl…
You Do Not Need More Data: Improving End-To-End Speech Recognition by Text-To-Speech Data Augmentation
Aleksandr Laptev, Roman Korostik, Aleksey Svischev +3
Data augmentation is one of the most effective ways to make end-to-end automatic speech recognition (ASR) perform close to the conventional hybrid approach, especially when dealing…
Towards a Competitive End-to-End Speech Recognition for CHiME-6 Dinner Party Transcription
Andrei Andrusenko, Aleksandr Laptev, Ivan Medennikov
While end-to-end ASR systems have proven competitive with the conventional hybrid approach, they are prone to accuracy degradation when it comes to noisy and low-resource condition…
Techniques for Vocabulary Expansion in Hybrid Speech Recognition Systems
Nikolay Malkovsky, Vladimir Bataev, Dmitrii Sviridkin +4
The problem of out of vocabulary words (OOV) is typical for any speech recognition system, hybrid systems are usually constructed to recognize a fixed set of words and rarely can i…