Cyclostationarity Analysis as a Complement to Self-Supervised Representations for Speech Deepfake Detection
arXiv:2603.03921
Abstract
Speech deepfake detection (SDD) is essential for maintaining trust in voice-driven technologies and digital media. Although recent SDD systems increasingly rely on SSL representations that capture rich contextual information, complementary signal-driven acoustic features remain important for modeling fine-grained structural properties of speech. Most existing acoustic front ends are based on time-frequency representations, which do not fully exploit higher-order spectral dependencies inherent in speech signals. We introduce a cyclostationarity-inspired acoustic feature extraction framework for SDD based on spectral correlation density (SCD). The proposed features model periodic statistical structures in speech by capturing spectral correlations between frequency components. In particular, we investigate whether cyclostationary representations provide complementary information beyond conventional time-frequency features and modern SSL embeddings. To this end, we introduce temporally structured SCD representations that preserve the evolution of spectral and cyclic-frequency correlations over time. Their effectiveness is evaluated using multiple SDD architectures, including conventional neural networks, SSL-based embedding systems, and hybrid fusion models. Experiments on ASVspoof 2019 LA, 2021 DF, and ASVspoof 5 demonstrate that higher-order spectral correlations provide complementary information to SSL embeddings in several challenging conditions. In particular, fusion of SSL and SCD embeddings reduces the equal error rate on ASVspoof 2019 LA from 8.28% to 0.98%, and from 15.81% to 14.80% on the challenging ASVspoof 5 dataset. These findings establish higher-order cyclostationary spectral correlations as a complementary source of information for modern SDD, while providing new signal-processing insight into the statistical differences between natural and synthetic speech.
accepted for publication in IEEE Transactions on Audio, Speech and Language Processing