23 citations · 34 across the 31 of their papers we have counts for
10 papers · 2 filters
The Spheres Dataset: Multitrack Orchestral Recordings for Music Source Separation and Information Retrieval
Jaime Garcia-Martinez, David Diaz-Guerra, John Anderson +5
This paper introduces The Spheres dataset, multitrack orchestral recordings designed to advance machine learning research in music source separation and related MIR tasks within th…
Acoustic Simulation Framework for Multi-channel Replay Speech Detection
Michael Neri, Tuomas Virtanen
Replay speech attacks pose a significant threat to voice-controlled systems, especially in smart environments where voice assistants are widely deployed. While multi-channel audio…
Multi-Utterance Speech Separation and Association Trained on Short Segments
Yuzhu Wang, Archontis Politis, Konstantinos Drossos +1
Current deep neural network (DNN) based speech separation faces a fundamental challenge -- while the models need to be trained on short segments due to computational constraints, r…
Inter-Speaker Relative Cues for Text-Guided Target Speech Extraction
Wang Dai, Archontis Politis, Tuomas Virtanen
We propose a novel approach that utilizes inter-speaker relative cues to distinguish target speakers and extract their voices from mixtures. Continuous cues (e.g., temporal order,…
Attractor-Based Speech Separation of Multiple Utterances by Unknown Number of Speakers
Yuzhu Wang, Archontis Politis, Konstantinos Drossos +1
This paper addresses the problem of single-channel speech separation, where the number of speakers is unknown, and each speaker may speak multiple utterances. We propose a speech s…
Impact of Microphone Array Mismatches to Learning-based Replay Speech Detection
Michael Neri, Tuomas Virtanen
In this work, we investigate the generalization of a multi-channel learning-based replay speech detector, which employs adaptive beamforming and detection, across different microph…