6 papers · 1 filter
Listen first: Output-based multi-microphone speech enhancement
Panos Apostolidis, Svend Feldt, Zheng-Hua Tan +2
Traditionally, hearing-aid speech enhancement (SE) algorithms rely on input-based feature estimation, often derived by a voice activity detector (VAD), to configure beamformers. Ye…
Ranking the Impact of Contextual Specialization in Neural Speech Enhancement
Peter Leer, Svend Feldt, Zheng-Hua Tan +2
We systematically investigate neural speech enhancement systems, ranging from very small (10\,k parameters) to medium-large (2-5\,M parameters), which specialize to aco…
Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining
Holger Severin Bovbjerg, Jan Ãstergaard, Jesper Jensen +1
Target-Speaker Voice Activity Detection (TS-VAD) is the task of detecting the presence of speech from a known target-speaker in an audio frame. Recently, deep neural network-based…
Hearing-Loss Compensation Using Deep Neural Networks: A Framework and Results From a Listening Test
Peter Leer, Jesper Jensen, Laurel H. Carney +3
This article investigates the use of deep neural networks (DNNs) for hearing-loss compensation. Hearing loss is a prevalent issue affecting millions of people worldwide, and conven…
Investigating the Design Space of Diffusion Models for Speech Enhancement
Philippe Gonzalez, Zheng-Hua Tan, Jan Ãstergaard +3
Diffusion models are a new class of generative models that have shown outstanding performance in image generation literature. As a consequence, studies have attempted to apply diff…
The Effect of Training Dataset Size on Discriminative and Diffusion-Based Speech Enhancement Systems
Philippe Gonzalez, Zheng-Hua Tan, Jan Ãstergaard +3
The performance of deep neural network-based speech enhancement systems typically increases with the training dataset size. However, studies that investigated the effect of trainin…