collaborators

5 papers

cs.SD2024

Audio-Visual Speaker Diarization: Current Databases, Approaches and Challenges

Victoria Mingote, Alfonso Ortega, Antonio Miguel +1

Nowadays, the large amount of audio-visual content available has fostered the need to develop new robust automatic speaker diarization systems to analyse and characterise it. This…

cs.LG2024

Predefined Prototypes for Intra-Class Separation and Disentanglement

Antonio Almudévar, Théo Mariotte, Alfonso Ortega +4

Prototypical Learning is based on the idea that there is a point (which we call prototype) around which the embeddings of a class are clustered. It has shown promising results in s…

eess.AS2024

Explainable by-design Audio Segmentation through Non-Negative Matrix Factorization and Probing

Martin Lebourdais, Théo Mariotte, Antonio Almudévar +2

Audio segmentation is a key task for many speech technologies, most of which are based on neural networks, usually considered as black boxes, with high-level performances. However,…

cs.LG2024

Unsupervised Multiple Domain Translation through Controlled Disentanglement in Variational Autoencoder

Antonio Almudévar, Théo Mariotte, Alfonso Ortega +1

Unsupervised Multiple Domain Translation is the task of transforming data from one domain to other domains without having paired data to train the systems. Typically, methods based…

eess.AS2024

An Explainable Proxy Model for Multiabel Audio Segmentation

Théo Mariotte, Antonio Almudévar, Marie Tahon +1

Audio signal segmentation is a key task for automatic audio indexing. It consists of detecting the boundaries of class-homogeneous segments in the signal. In many applications, exp…