3 papers
eess.AS2021
LT-LM: a novel non-autoregressive language model for single-shot lattice rescoring
Anton Mitrofanov, Mariya Korenevskaya, Ivan Podluzhny +7
Neural network-based language models are commonly used in rescoring approaches to improve the quality of modern automatic speech recognition (ASR) systems. Most of the existing met…
eess.AS2021
Dynamic Acoustic Unit Augmentation With BPE-Dropout for Low-Resource End-to-End Speech Recognition
Aleksandr Laptev, Andrei Andrusenko, Ivan Podluzhny +3
With the rapid development of speech assistants, adapting server-intended automatic speech recognition (ASR) solutions to a direct device has become crucial. Researchers and indust…
eess.AS2020
Target-Speaker Voice Activity Detection: a Novel Approach for Multi-Speaker Diarization in a Dinner Party Scenario
Ivan Medennikov, Maxim Korenevsky, Tatiana Prisyach +9
Speaker diarization for real-life scenarios is an extremely challenging problem. Widely used clustering-based diarization approaches perform rather poorly in such conditions, mainl…