5 papers
Audio-Visual Speaker Diarization: Current Databases, Approaches and Challenges
Victoria Mingote, Alfonso Ortega, Antonio Miguel +1
Nowadays, the large amount of audio-visual content available has fostered the need to develop new robust automatic speaker diarization systems to analyse and characterise it. This…
Predefined Prototypes for Intra-Class Separation and Disentanglement
Antonio Almudévar, Théo Mariotte, Alfonso Ortega +4
Prototypical Learning is based on the idea that there is a point (which we call prototype) around which the embeddings of a class are clustered. It has shown promising results in s…
Explainable by-design Audio Segmentation through Non-Negative Matrix Factorization and Probing
Martin Lebourdais, Théo Mariotte, Antonio Almudévar +2
Audio segmentation is a key task for many speech technologies, most of which are based on neural networks, usually considered as black boxes, with high-level performances. However,…
Unsupervised Multiple Domain Translation through Controlled Disentanglement in Variational Autoencoder
Antonio Almudévar, Théo Mariotte, Alfonso Ortega +1
Unsupervised Multiple Domain Translation is the task of transforming data from one domain to other domains without having paired data to train the systems. Typically, methods based…
An Explainable Proxy Model for Multiabel Audio Segmentation
Théo Mariotte, Antonio Almudévar, Marie Tahon +1
Audio signal segmentation is a key task for automatic audio indexing. It consists of detecting the boundaries of class-homogeneous segments in the signal. In many applications, exp…