A Hybrid Classical-Learning Framework for Adaptive Decision Directed Speech Enhancement
arXiv:2609.26183
Abstract
Speech enhancement aims to recover clean speech signals from noisy observations while preserving speech quality and intelligibility. Classical methods such as Spectral Subtraction and Decision-Directed (DD) enhancement remain widely used because of their interpretability and low computational complexity, but they may suffer from musical-noise artifacts or excessive attenuation of weak speech components under low signal-to-noise ratio (SNR) conditions. This paper proposes an Adaptive Beta-Constrained Decision-Directed (ABCDD) speech enhancement framework that extends the conventional DD method through a frame-dependent lower gain bound. The introduced beta parameter controls the tradeoff between noise suppression and speech preservation. To automate parameter selection for large and diverse datasets, a lightweight multilayer perceptron (MLP) model is further developed to predict frame-level beta values directly from noisy-speech features. The proposed framework is evaluated using both a representative speech example and large-scale testing on the VoiceBank-DEMAND dataset. In the representative example, ABCDD outperformed conventional Spectral Subtraction and classical DD across multiple objective metrics, including SNR, Log-Spectral Distance (LSD), Root-Mean-Square Error (RMSE), correlation, and Scale-Invariant Signal-to-Distortion Ratio (SI-SDR). On 100 unseen VoiceBank-DEMAND test files, the proposed MLP-beta ABCDD method improved average scale-aligned SNR from 9.41 dB to 13.82 dB, corresponding to an average gain of 4.41 dB. The results indicate that combining interpretable classical enhancement structure with lightweight machine-learning-based parameter adaptation provides an effective and practical direction for robust speech enhancement.
6 pages, 2 figures, 1 table