4 papers
Sparse Autoencoders Make Audio Foundation Models more Explainable
Théo Mariotte, Martin Lebourdais, Antonio Almudévar +3
Audio pretrained models are widely employed to solve various tasks in speech processing, sound event detection, or music information retrieval. However, the representations learned…
Predefined Prototypes for Intra-Class Separation and Disentanglement
Antonio Almudévar, Théo Mariotte, Alfonso Ortega +4
Prototypical Learning is based on the idea that there is a point (which we call prototype) around which the embeddings of a class are clustered. It has shown promising results in s…
Explainable by-design Audio Segmentation through Non-Negative Matrix Factorization and Probing
Martin Lebourdais, Théo Mariotte, Antonio Almudévar +2
Audio segmentation is a key task for many speech technologies, most of which are based on neural networks, usually considered as black boxes, with high-level performances. However,…
ASoBO: Attentive Beamformer Selection for Distant Speaker Diarization in Meetings
Theo Mariotte, Anthony Larcher, Silvio Montresor +1
Speaker Diarization (SD) aims at grouping speech segments that belong to the same speaker. This task is required in many speech-processing applications, such as rich meeting transc…