4 papers
DTT-BSR: GAN-based DTTNet with RoPE Transformer Enhancement for Music Source Restoration
Shihong Tan, Haoyu Wang, Youran Ni +8
Music source restoration (MSR) aims to recover unprocessed stems from mixed and mastered recordings. The challenge lies in both separating overlapping sources and reconstructing si…
Moving Speaker Separation via Parallel Spectral-Spatial Processing
Yuzhu Wang, Archontis Politis, Konstantinos Drossos +1
Multi-channel speech separation in dynamic environments is challenging as time-varying spatial and spectral features evolve at different temporal scales. Existing methods typically…
Multi-Utterance Speech Separation and Association Trained on Short Segments
Yuzhu Wang, Archontis Politis, Konstantinos Drossos +1
Current deep neural network (DNN) based speech separation faces a fundamental challenge -- while the models need to be trained on short segments due to computational constraints, r…
Attractor-Based Speech Separation of Multiple Utterances by Unknown Number of Speakers
Yuzhu Wang, Archontis Politis, Konstantinos Drossos +1
This paper addresses the problem of single-channel speech separation, where the number of speakers is unknown, and each speaker may speak multiple utterances. We propose a speech s…