4 papers
LT-LM: a novel non-autoregressive language model for single-shot lattice rescoring
Anton Mitrofanov, Mariya Korenevskaya, Ivan Podluzhny +7
Neural network-based language models are commonly used in rescoring approaches to improve the quality of modern automatic speech recognition (ASR) systems. Most of the existing met…
Target-Speaker Voice Activity Detection: a Novel Approach for Multi-Speaker Diarization in a Dinner Party Scenario
Ivan Medennikov, Maxim Korenevsky, Tatiana Prisyach +9
Speaker diarization for real-life scenarios is an extremely challenging problem. Widely used clustering-based diarization approaches perform rather poorly in such conditions, mainl…
Exploring Gaussian mixture model framework for speaker adaptation of deep neural network acoustic models
Natalia Tomashenko, Yuri Khokhlov, Yannick Esteve
In this paper we investigate the GMM-derived (GMMD) features for adaptation of deep neural network (DNN) acoustic models. The adaptation of the DNN trained on GMMD features is done…
Fast and Accurate OOV Decoder on High-Level Features
Yuri Khokhlov, Natalia Tomashenko, Ivan Medennikov +1
This work proposes a novel approach to out-of-vocabulary (OOV) keyword search (KWS) task. The proposed approach is based on using high-level features from an automatic speech recog…