3 papers
cs.SD2026
Masked Autoregressive Speech Enhancement with Continuous Neural Audio Codec Representations
Yoto Fujita, Simon Leglaive, Laurent Girin
Most previous work on speech enhancement (SE) based on masked generative modeling relied on discrete token representations of audio signals, obtained using neural audio codecs (NAC…
cs.SD2024
Run-Time Adaptation of Neural Beamforming for Robust Speech Dereverberation and Denoising
Yoto Fujita, Aditya Arie Nugraha, Diego Di Carlo +3
This paper describes speech enhancement for realtime automatic speech recognition (ASR) in real environments. A standard approach to this task is to use neural beamforming that can…
cs.SD2024
DOA-Aware Audio-Visual Self-Supervised Learning for Sound Event Localization and Detection
Yoto Fujita, Yoshiaki Bando, Keisuke Imoto +2
This paper describes sound event localization and detection (SELD) for spatial audio recordings captured by firstorder ambisonics (FOA) microphones. In this task, one may train a d…