4 papers
Language Modelling for Speaker Diarization in Telephonic Interviews
Miquel India, Javier Hernando, José A. R. Fonollosa
The aim of this paper is to investigate the benefit of combining both language and acoustic modelling for speaker diarization. Although conventional systems only use acoustic featu…
BSC-UPC at EmoSPeech-IberLEF2024: Attention Pooling for Emotion Recognition
Marc Casals-Salvador, Federico Costa, Miquel India +1
The domain of speech emotion recognition (SER) has persistently been a frontier within the landscape of machine learning. It is an active field that has been revolutionized in the…
Double Multi-Head Attention Multimodal System for Odyssey 2024 Speech Emotion Recognition Challenge
Federico Costa, Miquel India, Javier Hernando
As computer-based applications are becoming more integrated into our daily lives, the importance of Speech Emotion Recognition (SER) has increased significantly. Promoting research…
Speaker Characterization by means of Attention Pooling
Federico Costa, Miquel India, Javier Hernando
State-of-the-art Deep Learning systems for speaker verification are commonly based on speaker embedding extractors. These architectures are usually composed of a feature extractor…