activity
20232026
most citedSpeaker Distance Estimation in Enclosures from Single-Channel Audio

23 citations · 34 across the 31 of their papers we have counts for

collaborators
Showing 2025 · eess.ASShow all

10 papers · 2 filters

eess.AS2025

The Spheres Dataset: Multitrack Orchestral Recordings for Music Source Separation and Information Retrieval

Jaime Garcia-Martinez, David Diaz-Guerra, John Anderson +5

This paper introduces The Spheres dataset, multitrack orchestral recordings designed to advance machine learning research in music source separation and related MIR tasks within th…

eess.AS2025

Acoustic Simulation Framework for Multi-channel Replay Speech Detection

Michael Neri, Tuomas Virtanen

Replay speech attacks pose a significant threat to voice-controlled systems, especially in smart environments where voice assistants are widely deployed. While multi-channel audio…

eess.AS2025

Multi-Utterance Speech Separation and Association Trained on Short Segments

Yuzhu Wang, Archontis Politis, Konstantinos Drossos +1

Current deep neural network (DNN) based speech separation faces a fundamental challenge -- while the models need to be trained on short segments due to computational constraints, r…

eess.AS2025

Inter-Speaker Relative Cues for Text-Guided Target Speech Extraction

Wang Dai, Archontis Politis, Tuomas Virtanen

We propose a novel approach that utilizes inter-speaker relative cues to distinguish target speakers and extract their voices from mixtures. Continuous cues (e.g., temporal order,…

eess.AS2025

Attractor-Based Speech Separation of Multiple Utterances by Unknown Number of Speakers

Yuzhu Wang, Archontis Politis, Konstantinos Drossos +1

This paper addresses the problem of single-channel speech separation, where the number of speakers is unknown, and each speaker may speak multiple utterances. We propose a speech s…

eess.AS2025★ 1 cited

Impact of Microphone Array Mismatches to Learning-based Replay Speech Detection

Michael Neri, Tuomas Virtanen

In this work, we investigate the generalization of a multi-channel learning-based replay speech detector, which employs adaptive beamforming and detection, across different microph…