Configurable EBEN: Extreme Bandwidth Extension Network to enhance body-conducted speech capture
arXiv:2303.10008 · doi:10.1109/TASLP.2023.3313433
Abstract
This paper presents a configurable version of Extreme Bandwidth Extension Network (EBEN), a Generative Adversarial Network (GAN) designed to improve audio captured with body-conduction microphones. We show that although these microphones significantly reduce environmental noise, this insensitivity to ambient noise happens at the expense of the bandwidth of the speech signal acquired by the wearer of the devices. The obtained captured signals therefore require the use of signal enhancement techniques to recover the full-bandwidth speech. EBEN leverages a configurable multiband decomposition of the raw captured signal. This decomposition allows the data time domain dimensions to be reduced and the full band signal to be better controlled. The multiband representation of the captured signal is processed through a U-Net-like model, which combines feature and adversarial losses to generate an enhanced speech signal. We also benefit from this original representation in the proposed configurable discriminators architecture. The configurable EBEN approach can achieve state-of-the-art enhancement results on synthetic data with a lightweight generator that allows real-time processing.
Accepted in IEEE/ACM Transactions on Audio, Speech and Language Processing on 14/08/2023
References in corpus (5)
- MLS: A Large-Scale Multilingual Dataset for Speech Research
- High Fidelity Neural Audio Compression
- Waveform Modeling and Generation Using Hierarchical Recurrent Neural Networks for Speech Bandwidth Extension
- EBEN: Extreme bandwidth extension network applied to speech signals captured with noise-resilient body-conduction microphones
- Training Strategies for Own Voice Reconstruction in Hearing Protection Devices using an In-ear Microphone
Cited by in corpus (4)
- Modeling of Speech-dependent Own Voice Transfer Characteristics for Hearables with In-ear Microphones
- Throat and acoustic paired speech dataset for deep learning-based speech enhancement
- Low-Complexity Own Voice Reconstruction for Hearables with an In-Ear Microphone
- Speech-dependent Data Augmentation for Own Voice Reconstruction with Hearable Microphones in Noisy Environments