collaborators

5 papers

eess.AS2026

Temporal Pooling Strategies for Training-Free Anomalous Sound Detection with Self-Supervised Audio Embeddings

Kevin Wilkinghoff, Sarthak Yadav, Zheng-Hua Tan

Training-free anomalous sound detection (ASD) based on pre-trained audio embedding models has recently garnered significant attention, as it enables the detection of anomalous soun…

cs.SD2025

An overview of neural architectures for self-supervised audio representation learning from masked spectrograms

Sarthak Yadav, Sergios Theodoridis, Zheng-Hua Tan

In recent years, self-supervised learning has amassed significant interest for training deep neural representations without labeled data. One such self-supervised learning approach…

cs.SD2025

AxLSTMs: learning self-supervised audio representations with xLSTMs

Sarthak Yadav, Sergios Theodoridis, Zheng-Hua Tan

While the transformer has emerged as the eminent neural architecture, several independent lines of research have emerged to address its limitations. Recurrent neural approaches hav…

cs.SD2025

AudioMAE++: learning better masked audio representations with SwiGLU FFNs

Sarthak Yadav, Sergios Theodoridis, Zheng-Hua Tan

Masked Autoencoders (MAEs) trained on audio spectrogram patches have emerged as a prominent approach for learning self-supervised audio representations. While several recent papers…

cs.SD2024

Audio Mamba: Selective State Spaces for Self-Supervised Audio Representations

Sarthak Yadav, Zheng-Hua Tan

Despite its widespread adoption as the prominent neural architecture, the Transformer has spurred several independent lines of work to address its limitations. One such approach is…