paper

cSTMM: A Unified Complex Spherical Student's Mixture Model for Directional Statistics in Mask-Based Blind Speech Separation

arXiv:2605.25512

Abstract

Directional-statistics-based mask-based blind speech separation (BSS) clusters normalized time-frequency (TF) observations from microphones on the complex unit sphere, without relying on plane-wave or spherical-wave assumptions. Existing methods use separately defined angular mixture models, which makes the effect of density shape difficult to isolate. This paper proposes a complex spherical Student's mixture model (cSTMM) that connects the complex angular central Gaussian mixture model (cACGMM), complex Bingham mixture model (cBMM), and complex Watson mixture model (cWMM) through the degrees of freedom and eigenvalue constraints. We derive a latent-scale expectation-maximization (EM) framework with an approximate M-step based on high-concentration approximation (HCA). On noise-free LibriSpeech mixtures reverberated using measured room impulse responses (RIRs), the development-selected value outperformed the cACGMM-equivalent choice in all 18 test conditions, yielding a mean signal-to-distortion ratio improvement (SDRi) gain of . The model reduces to the cACGMM at and approaches the cBMM/cWMM in the large- limits.

cSTMM: A Unified Complex Spherical Student's $t$ Mixture Model for Directional Statistics in Mask-Based Blind Speech Separation · wovepaper