activity
20192022
most citedStreaming Multi-speaker ASR with RNN-T

3 citations · 3 across the 3 of their papers we have counts for

collaborators

5 papers

eess.AS2022

Separator-Transducer-Segmenter: Streaming Recognition and Segmentation of Multi-party Speech

Ilya Sklyar, Anna Piunova, Christian Osendorfer

Streaming recognition and segmentation of multi-party conversations with overlapping speech is crucial for the next generation of voice assistant applications. In this work we addr…

eess.AS20203 cited

Streaming Multi-speaker ASR with RNN-T

Ilya Sklyar, Anna Piunova, Yulan Liu

Recent research shows end-to-end ASR systems can recognize overlapped speech from multiple speakers. However, all published works have assumed no latency constraints during inferen…

eess.AS2020

Improving RNN-T ASR Accuracy Using Context Audio

Andreas Schwarz, Ilya Sklyar, Simon Wiesler

We present a training scheme for streaming automatic speech recognition (ASR) based on recurrent neural network transducers (RNN-T) which allows the encoder network to learn to exp…

eess.AS2020

Subword Regularization: An Analysis of Scalability and Generalization for End-to-End Automatic Speech Recognition

Egor Lakomkin, Jahn Heymann, Ilya Sklyar +1

Subwords are the most widely used output units in end-to-end speech recognition. They combine the best of two worlds by modeling the majority of frequent words directly and at the…

cs.SD2019

Analysis of Deep Clustering as Preprocessing for Automatic Speech Recognition of Sparsely Overlapping Speech

Tobias Menne, Ilya Sklyar, Ralf Schlüter +1

Significant performance degradation of automatic speech recognition (ASR) systems is observed when the audio signal contains cross-talk. One of the recently proposed approaches to…