activity
20242026
most citedHierarchical Label Propagation: A Model-Size-Dependent Performance Booster for AudioSet Tagging

1 citations · 1 across the 3 of their papers we have counts for

collaborators
Showing cs.SDShow all

8 papers · 1 filter

cs.SD2026

BEST-RQ-2: Contextualize-Then-Predict, a Two-Step Approach for Self-Supervised Audio Representations

Ludovic K. Tuncay, Etienne Labbé, Thomas Pellegrini

Self-supervised learning enables audio representations that transfer across domains and tasks. We present BEST-RQ-2, an evolution of BEST-RQ that retains frozen randomprojection-ba…

cs.SD2025

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning

Ludovic Tuncay, Etienne Labbé, Emmanouil Benetos +1

Building on the Joint-Embedding Predictive Architecture (JEPA) paradigm, a recent self-supervised learning framework that predicts latent representations of masked regions in high-…

cs.SD20251 cited

Hierarchical Label Propagation: A Model-Size-Dependent Performance Booster for AudioSet Tagging

Ludovic Tuncay, Etienne Labbé, Thomas Pellegrini

AudioSet is one of the most used and largest datasets in audio tagging, containing about 2 million audio samples that are manually labeled with 527 event categories organized into…

cs.SD20254 cited

Self-Supervised Models for Phoneme Recognition: Applications in Children's Speech for Reading Learning

Lucas Block Medin, Thomas Pellegrini, Lucile Gelin

Child speech recognition is still an underdeveloped area of research due to the lack of data (especially on non-English languages) and the specific difficulties of this task. Havin…

cs.SD2024

Enhancing Synthetic Training Data for Speech Commands: From ASR-Based Filtering to Domain Adaptation in SSL Latent Space

Sebastião Quintas, Isabelle Ferrané, Thomas Pellegrini

The use of synthetic speech as data augmentation is gaining increasing popularity in fields such as automatic speech recognition and speech classification tasks. Despite novel text…

cs.SD2023

Multilingual Audio Captioning using machine translated data

Matéo Cousin, Étienne Labbé, Thomas Pellegrini

Automated Audio Captioning (AAC) systems attempt to generate a natural language sentence, a caption, that describes the content of an audio recording, in terms of sound events. Exi…