8 papers
Listen first: Output-based multi-microphone speech enhancement
Panos Apostolidis, Svend Feldt, Zheng-Hua Tan +2
Traditionally, hearing-aid speech enhancement (SE) algorithms rely on input-based feature estimation, often derived by a voice activity detector (VAD), to configure beamformers. Ye…
Ranking the Impact of Contextual Specialization in Neural Speech Enhancement
Peter Leer, Svend Feldt, Zheng-Hua Tan +2
We systematically investigate neural speech enhancement systems, ranging from very small (10\,k parameters) to medium-large (2-5\,M parameters), which specialize to aco…
Exploring Resolution-Wise Shared Attention in Hybrid Mamba-U-Nets for Improved Cross-Corpus Speech Enhancement
Nikolai Lund Kühne, Jesper Jensen, Jan Ãstergaard +1
Recent advances in speech enhancement have shown that models combining Mamba and attention mechanisms yield superior cross-corpus generalization performance. At the same time, inte…
MambAttention: Mamba with Multi-Head Attention for Generalizable Single-Channel Speech Enhancement
Nikolai Lund Kühne, Jesper Jensen, Jan Ãstergaard +1
With new sequence models like Mamba and xLSTM, several studies have shown that these models match or outperform the state-of-the-art in single-channel speech enhancement and audio…
Learning Robust Spatial Representations from Binaural Audio through Feature Distillation
Holger Severin Bovbjerg, Jan Ãstergaard, Jesper Jensen +2
Recently, deep representation learning has shown strong performance in multiple audio tasks. However, its use for learning spatial representations from multichannel audio is undere…
xLSTM-SENet: xLSTM for Single-Channel Speech Enhancement
Nikolai Lund Kühne, Jan Ãstergaard, Jesper Jensen +1
While attention-based architectures, such as Conformers, excel in speech enhancement, they face challenges such as scalability with respect to input sequence length. In contrast, t…