2 papers
cs.SD2025
Learning Robust Spatial Representations from Binaural Audio through Feature Distillation
Holger Severin Bovbjerg, Jan Ãstergaard, Jesper Jensen +2
Recently, deep representation learning has shown strong performance in multiple audio tasks. However, its use for learning spatial representations from multichannel audio is undere…
eess.AS2025
Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining
Holger Severin Bovbjerg, Jan Ãstergaard, Jesper Jensen +1
Target-Speaker Voice Activity Detection (TS-VAD) is the task of detecting the presence of speech from a known target-speaker in an audio frame. Recently, deep neural network-based…