12 citations · 19 across the 14 of their papers we have counts for
8 papers · 1 filter
Retrieval-Augmented Neural Field for HRTF Upsampling and Personalization
Yoshiki Masuyama, Gordon Wichern, François G. Germain +2
Head-related transfer functions (HRTFs) with dense spatial grids are desired for immersive binaural audio generation, but their recording is time-consuming. Although HRTF spatial u…
Leveraging Audio-Only Data for Text-Queried Target Sound Extraction
Kohei Saijo, Janek Ebbers, François G. Germain +3
The goal of text-queried target sound extraction (TSE) is to extract from a mixture a sound source specified with a natural-language caption. While it is preferable to have access…
Reverberation as Supervision for Speech Separation
Rohith Aralikatti, Christoph Boeddeker, Gordon Wichern +2
This paper proposes reverberation as supervision (RAS), a novel unsupervised loss function for single-channel reverberant speech separation. Prior methods for unsupervised separati…
Locate This, Not That: Class-Conditioned Sound Event DOA Estimation
Olga Slizovskaia, Gordon Wichern, Zhong-Qiu Wang +1
Existing systems for sound event localization and detection (SELD) typically operate by estimating a source location for all classes at every time instant. In this paper, we propos…
Advancing Momentum Pseudo-Labeling with Conformer and Initialization Strategy
Yosuke Higuchi, Niko Moritz, Jonathan Le Roux +1
Pseudo-labeling (PL), a semi-supervised learning (SSL) method where a seed model performs self-training using pseudo-labels generated from untranscribed speech, has been shown to e…
Dual Causal/Non-Causal Self-Attention for Streaming End-to-End Speech Recognition
Niko Moritz, Takaaki Hori, Jonathan Le Roux
Attention-based end-to-end automatic speech recognition (ASR) systems have recently demonstrated state-of-the-art results for numerous tasks. However, the application of self-atten…