4 papers
Exploring Resolution-Wise Shared Attention in Hybrid Mamba-U-Nets for Improved Cross-Corpus Speech Enhancement
Nikolai Lund Kühne, Jesper Jensen, Jan Østergaard +1
Recent advances in speech enhancement have shown that models combining Mamba and attention mechanisms yield superior cross-corpus generalization performance. At the same time, inte…
Learning Robust Spatial Representations from Binaural Audio through Feature Distillation
Holger Severin Bovbjerg, Jan Østergaard, Jesper Jensen +2
Recently, deep representation learning has shown strong performance in multiple audio tasks. However, its use for learning spatial representations from multichannel audio is undere…
Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining
Holger Severin Bovbjerg, Jan Østergaard, Jesper Jensen +1
Target-Speaker Voice Activity Detection (TS-VAD) is the task of detecting the presence of speech from a known target-speaker in an audio frame. Recently, deep neural network-based…
xLSTM-SENet: xLSTM for Single-Channel Speech Enhancement
Nikolai Lund Kühne, Jan Østergaard, Jesper Jensen +1
While attention-based architectures, such as Conformers, excel in speech enhancement, they face challenges such as scalability with respect to input sequence length. In contrast, t…