3 papers
cs.SD2026
BEST-RQ-2: Contextualize-Then-Predict, a Two-Step Approach for Self-Supervised Audio Representations
Ludovic K. Tuncay, Etienne Labbé, Thomas Pellegrini
Self-supervised learning enables audio representations that transfer across domains and tasks. We present BEST-RQ-2, an evolution of BEST-RQ that retains frozen randomprojection-ba…
cs.SD2025
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning
Ludovic Tuncay, Etienne Labbé, Emmanouil Benetos +1
Building on the Joint-Embedding Predictive Architecture (JEPA) paradigm, a recent self-supervised learning framework that predicts latent representations of masked regions in high-…
cs.SD2025
Hierarchical Label Propagation: A Model-Size-Dependent Performance Booster for AudioSet Tagging
Ludovic Tuncay, Etienne Labbé, Thomas Pellegrini
AudioSet is one of the most used and largest datasets in audio tagging, containing about 2 million audio samples that are manually labeled with 527 event categories organized into…