UD-ASD: A Unified Diffusion Model for Anomalous Sound Detection
arXiv:2607.12576
The paper introduces a unified diffusion model that uses machine ID conditioning to reconstruct log‑Mel spectrograms for anomalous sound detection across multiple machine types, improving detection metrics on the DCASE2022 benchmark.
Abstract
Anomalous Sound Detection (ASD) aims to determine whether faults have occurred by monitoring sounds. Existing methods detect a limited range of anomalies, exhibit poor generalization, or train a separate model for each machine. Diffusion models possess strong generalization and can generate specific data with condition guidance. We propose a unified diffusion model only with a small module. The audio is first transformed into log-Mel spectrograms. The lightweight module embeds machine IDs into condition embeddings, guiding the model to reconstruct data for specific machines. Then diffusion model reconstructs data with condition, using Gaussian Mixture Models to fit the distributions of reconstruction errors. Our unified model could monitor multiple machine types and learn more fundamental feature spaces with cross-domain learning. Experiments on DCASE2022 Challenge Task 2 show that our model achieves 3.44% AUC and 2.52% pAUC improvements over baseline, validating its effectiveness.
5 pages, 3 figures, Interspeech 2026