46 citations · 80 across the 17 of their papers we have counts for
14 papers
Mask-based Neural Beamforming for Moving Speakers with Self-Attention-based Tracking
Tsubasa Ochiai, Marc Delcroix, Tomohiro Nakatani +1
Beamforming is a powerful tool designed to enhance speech signals from the direction of a target source. Computing the beamforming filter requires estimating spatial covariance mat…
Few-shot learning of new sound classes for target sound extraction
Marc Delcroix, Jorge Bennasar Vázquez, Tsubasa Ochiai +2
Target sound extraction consists of extracting the sound of a target acoustic event (AE) class from a mixture of AE sounds. It can be realized using a neural network that extracts…
PILOT: Introducing Transformers for Probabilistic Sound Event Localization
Christopher Schymura, Benedikt Bönninghoff, Tsubasa Ochiai +5
Sound event localization aims at estimating the positions of sound sources in the environment with respect to an acoustic receiver (e.g. a microphone array). Recent advances in thi…
Exploiting Attention-based Sequence-to-Sequence Architectures for Sound Event Localization
Christopher Schymura, Tsubasa Ochiai, Marc Delcroix +4
Sound event localization frameworks based on deep neural networks have shown increased robustness with respect to reverberation and noise in comparison to classical parametric appr…
Data Fusion for Audiovisual Speaker Localization: Extending Dynamic Stream Weights to the Spatial Domain
Julio Wissing, Benedikt Boenninghoff, Dorothea Kolossa +6
Estimating the positions of multiple speakers can be helpful for tasks like automatic speech recognition or speaker diarization. Both applications benefit from a known speaker posi…
Speaker activity driven neural speech extraction
Marc Delcroix, Katerina Zmolikova, Tsubasa Ochiai +2
Target speech extraction, which extracts the speech of a target speaker in a mixture given auxiliary speaker clues, has recently received increased interest. Various clues have bee…