1 citations · 1 across the 3 of their papers we have counts for
8 papers · 1 filter
BEST-RQ-2: Contextualize-Then-Predict, a Two-Step Approach for Self-Supervised Audio Representations
Ludovic K. Tuncay, Etienne Labbé, Thomas Pellegrini
Self-supervised learning enables audio representations that transfer across domains and tasks. We present BEST-RQ-2, an evolution of BEST-RQ that retains frozen randomprojection-ba…
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning
Ludovic Tuncay, Etienne Labbé, Emmanouil Benetos +1
Building on the Joint-Embedding Predictive Architecture (JEPA) paradigm, a recent self-supervised learning framework that predicts latent representations of masked regions in high-…
Hierarchical Label Propagation: A Model-Size-Dependent Performance Booster for AudioSet Tagging
Ludovic Tuncay, Etienne Labbé, Thomas Pellegrini
AudioSet is one of the most used and largest datasets in audio tagging, containing about 2 million audio samples that are manually labeled with 527 event categories organized into…
Self-Supervised Models for Phoneme Recognition: Applications in Children's Speech for Reading Learning
Lucas Block Medin, Thomas Pellegrini, Lucile Gelin
Child speech recognition is still an underdeveloped area of research due to the lack of data (especially on non-English languages) and the specific difficulties of this task. Havin…
Enhancing Synthetic Training Data for Speech Commands: From ASR-Based Filtering to Domain Adaptation in SSL Latent Space
Sebastião Quintas, Isabelle Ferrané, Thomas Pellegrini
The use of synthetic speech as data augmentation is gaining increasing popularity in fields such as automatic speech recognition and speech classification tasks. Despite novel text…
Multilingual Audio Captioning using machine translated data
Matéo Cousin, Étienne Labbé, Thomas Pellegrini
Automated Audio Captioning (AAC) systems attempt to generate a natural language sentence, a caption, that describes the content of an audio recording, in terms of sound events. Exi…