8 papers
Exploring Resolution-Wise Shared Attention in Hybrid Mamba-U-Nets for Improved Cross-Corpus Speech Enhancement
Nikolai Lund Kühne, Jesper Jensen, Jan Ãstergaard +1
Recent advances in speech enhancement have shown that models combining Mamba and attention mechanisms yield superior cross-corpus generalization performance. At the same time, inte…
MambAttention: Mamba with Multi-Head Attention for Generalizable Single-Channel Speech Enhancement
Nikolai Lund Kühne, Jesper Jensen, Jan Ãstergaard +1
With new sequence models like Mamba and xLSTM, several studies have shown that these models match or outperform the state-of-the-art in single-channel speech enhancement and audio…
A Steered Response Power Method for Sound Source Localization With Generic Acoustic Models
Kaspar Müller, Markus Buck, Simon Doclo +2
The steered response power (SRP) method is one of the most popular approaches for acoustic source localization with microphone arrays. It is often based on simplifying acoustic ass…
Learning Robust Spatial Representations from Binaural Audio through Feature Distillation
Holger Severin Bovbjerg, Jan Ãstergaard, Jesper Jensen +2
Recently, deep representation learning has shown strong performance in multiple audio tasks. However, its use for learning spatial representations from multichannel audio is undere…
Head-steered channel selection method for hearing aid applications using remote microphones
Vasudha Sathyapriyan, Michael S. Pedersen, Mike Brookes +3
We propose a channel selection method for hearing aid applications using remote microphones, in the presence of multiple competing talkers. The proposed channel selection method us…
xLSTM-SENet: xLSTM for Single-Channel Speech Enhancement
Nikolai Lund Kühne, Jan Ãstergaard, Jesper Jensen +1
While attention-based architectures, such as Conformers, excel in speech enhancement, they face challenges such as scalability with respect to input sequence length. In contrast, t…