6 citations · 30 across the 24 of their papers we have counts for
4 papers · 2 filters
MultiSV: Dataset for Far-Field Multi-Channel Speaker Verification
Ladislav Mošner, Oldřich Plchot, Lukáš Burget +1
Motivated by unconsolidated data situation and the lack of a standard benchmark in the field, we complement our previous efforts and present a comprehensive corpus designed for tra…
Revisiting joint decoding based multi-talker speech recognition with DNN acoustic model
Martin Kocour, Kateřina Žmolíková, Lucas Ondel +5
In typical multi-talker speech recognition systems, a neural network-based acoustic model predicts senone state posteriors for each speaker. These are later used by a single-talker…
EAT: Enhanced ASR-TTS for Self-supervised Speech Recognition
Murali Karthick Baskar, Lukáš Burget, Shinji Watanabe +2
Self-supervised ASR-TTS models suffer in out-of-domain data conditions. Here we propose an enhanced ASR-TTS (EAT) model that incorporates two main features: 1) The ASR…
Speaker embeddings by modeling channel-wise correlations
Themos Stafylakis, Johan Rohdin, Lukas Burget
Speaker embeddings extracted with deep 2D convolutional neural networks are typically modeled as projections of first and second order statistics of channel-frequency pairs onto a…