activity
20232026
most citedInvestigating the Design Space of Diffusion Models for Speech Enhancement

19 citations · 40 across the 13 of their papers we have counts for

collaborators

14 papers

eess.AS2026

Listen first: Output-based multi-microphone speech enhancement

Panos Apostolidis, Svend Feldt, Zheng-Hua Tan +2

Traditionally, hearing-aid speech enhancement (SE) algorithms rely on input-based feature estimation, often derived by a voice activity detector (VAD), to configure beamformers. Ye…

eess.AS2026

Ranking the Impact of Contextual Specialization in Neural Speech Enhancement

Peter Leer, Svend Feldt, Zheng-Hua Tan +2

We systematically investigate neural speech enhancement systems, ranging from very small (10\,k parameters) to medium-large (2-5\,M parameters), which specialize to aco…

cs.SD2025

Exploring Resolution-Wise Shared Attention in Hybrid Mamba-U-Nets for Improved Cross-Corpus Speech Enhancement

Nikolai Lund Kühne, Jesper Jensen, Jan Østergaard +1

Recent advances in speech enhancement have shown that models combining Mamba and attention mechanisms yield superior cross-corpus generalization performance. At the same time, inte…

cs.SD2025

Learning Robust Spatial Representations from Binaural Audio through Feature Distillation

Holger Severin Bovbjerg, Jan Østergaard, Jesper Jensen +2

Recently, deep representation learning has shown strong performance in multiple audio tasks. However, its use for learning spatial representations from multichannel audio is undere…

cs.SD2025★ 3 cited

MambAttention: Mamba with Multi-Head Attention for Generalizable Single-Channel Speech Enhancement

Nikolai Lund Kühne, Jesper Jensen, Jan Østergaard +1

With new sequence models like Mamba and xLSTM, several studies have shown that these models match or outperform the state-of-the-art in single-channel speech enhancement and audio…

eess.AS2025

Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining

Holger Severin Bovbjerg, Jan Østergaard, Jesper Jensen +1

Target-Speaker Voice Activity Detection (TS-VAD) is the task of detecting the presence of speech from a known target-speaker in an audio frame. Recently, deep neural network-based…