Unfolded Recursive Expectation-Maximization Neural Network For Speaker Tracking
arXiv:2607.26575
The paper introduces a deep‑unfolded recursive expectation‑maximization (REM) neural network that learns adaptive step‑size updates for tracking a single moving speaker in mildly reverberant rooms, achieving lower RMSE than a classical CREM baseline.
Abstract
We propose a deep unfolded REM network for robust tracking of a single moving speaker in mild reverberant environments. Unlike classical REM algorithms, which rely on fixed-step-size decay schedules, the proposed architecture learns an adaptive update policy by unfolding the iterative procedure into differentiable layers. We introduce a Step Size Network that leverages FiLM and PE to dynamically adjust the recursion weights based on temporal context and convergence state. Experimental results for tracking a single speaker under reverberant conditions demonstrate that the proposed unfolded network outperforms the classical CREM baseline, which employs a spatial grid search to map the estimated centroids to physical positions. In the single-speaker tracking task, the proposed method achieves a lower RMSE than the CREM baseline, highlighting its potential for dynamic acoustic scenarios.
proceedings of IWAENC 2026