4 papers
Temporal Pooling Strategies for Training-Free Anomalous Sound Detection with Self-Supervised Audio Embeddings
Kevin Wilkinghoff, Sarthak Yadav, Zheng-Hua Tan
Training-free anomalous sound detection (ASD) based on pre-trained audio embedding models has recently garnered significant attention, as it enables the detection of anomalous soun…
An overview of neural architectures for self-supervised audio representation learning from masked spectrograms
Sarthak Yadav, Sergios Theodoridis, Zheng-Hua Tan
In recent years, self-supervised learning has amassed significant interest for training deep neural representations without labeled data. One such self-supervised learning approach…
AxLSTMs: learning self-supervised audio representations with xLSTMs
Sarthak Yadav, Sergios Theodoridis, Zheng-Hua Tan
While the transformer has emerged as the eminent neural architecture, several independent lines of research have emerged to address its limitations. Recurrent neural approaches hav…
AudioMAE++: learning better masked audio representations with SwiGLU FFNs
Sarthak Yadav, Sergios Theodoridis, Zheng-Hua Tan
Masked Autoencoders (MAEs) trained on audio spectrogram patches have emerged as a prominent approach for learning self-supervised audio representations. While several recent papers…