1 citations · 1 across the 5 of their papers we have counts for
5 papers
An overview of neural architectures for self-supervised audio representation learning from masked spectrograms
Sarthak Yadav, Sergios Theodoridis, Zheng-Hua Tan
In recent years, self-supervised learning has amassed significant interest for training deep neural representations without labeled data. One such self-supervised learning approach…
AudioMAE++: learning better masked audio representations with SwiGLU FFNs
Sarthak Yadav, Sergios Theodoridis, Zheng-Hua Tan
Masked Autoencoders (MAEs) trained on audio spectrogram patches have emerged as a prominent approach for learning self-supervised audio representations. While several recent papers…
BiSSL: Enhancing the Alignment Between Self-Supervised Pretraining and Downstream Fine-Tuning via Bilevel Optimization
Gustav Wagner Zakarias, Lars Kai Hansen, Zheng-Hua Tan
Models initialized from self-supervised pretraining may suffer from poor alignment with downstream tasks, reducing the extent to which subsequent fine-tuning can adapt pretrained f…
AxLSTMs: learning self-supervised audio representations with xLSTMs
Sarthak Yadav, Sergios Theodoridis, Zheng-Hua Tan
While the transformer has emerged as the eminent neural architecture, several independent lines of research have emerged to address its limitations. Recurrent neural approaches hav…
Audio Mamba: Selective State Spaces for Self-Supervised Audio Representations
Sarthak Yadav, Zheng-Hua Tan
Despite its widespread adoption as the prominent neural architecture, the Transformer has spurred several independent lines of work to address its limitations. One such approach is…