activity
20182024
most citedAcoustic Scene Classification Using Fusion of Attentive Convolutional Neural Networks for DCASE2019 Challenge

6 citations · 16 across the 10 of their papers we have counts for

collaborators
Showing eess.ASShow all

16 papers · 1 filter

eess.AS2024

DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition

Alexander Polok, Dominik Klement, Martin Kocour +7

Speaker-attributed automatic speech recognition (ASR) in multi-speaker environments remains a significant challenge, particularly when systems conditioned on speaker embeddings fai…

eess.AS20221 cited

Extracting speaker and emotion information from self-supervised speech models via channel-wise correlations

Themos Stafylakis, Ladislav Mosner, Sofoklis Kakouros +3

Self-supervised learning of speech representations from large amounts of unlabeled data has enabled state-of-the-art results in several speech processing tasks. Aggregating these s…

eess.AS20221 cited

An attention-based backend allowing efficient fine-tuning of transformer models for speaker verification

Junyi Peng, Oldrich Plchot, Themos Stafylakis +3

In recent years, self-supervised learning paradigm has received extensive attention due to its great success in various down-stream tasks. However, the fine-tuning strategies for a…

eess.AS2022

DPCCN: Densely-Connected Pyramid Complex Convolutional Network for Robust Speech Separation And Extraction

Jiangyu Han, Yanhua Long, Lukas Burget +1

In recent years, a number of time-domain speech separation methods have been proposed. However, most of them are very sensitive to the environments and wide domain coverage tasks.…

eess.AS2021

EAT: Enhanced ASR-TTS for Self-supervised Speech Recognition

Murali Karthick Baskar, Lukáš Burget, Shinji Watanabe +2

Self-supervised ASR-TTS models suffer in out-of-domain data conditions. Here we propose an enhanced ASR-TTS (EAT) model that incorporates two main features: 1) The ASR

eess.AS2020

Integration of variational autoencoder and spatial clustering for adaptive multi-channel neural speech separation

Katerina Zmolikova, Marc Delcroix, Lukáš Burget +2

In this paper, we propose a method combining variational autoencoder model of speech with a spatial clustering approach for multi-channel speech separation. The advantage of integr…