6 papers
Speaker Group Encoding in Self-supervised Speech Recognition Models
Felix Herron, Solange Rossato Alexandre Allauzen, Benoit Favre +1
We investigate what self-supervised speech recognition models (S3Ms) learn about speaker groups (SGs). We examine several states of S3Ms: pretrained, finetuned on speaker identific…
Responsible Benchmarking of Fairness for Automatic Speech Recognition
Felix Herron, Ange Richard, François Portet +2
Many studies have shown automatic speech processing (ASR) systems have unequal performance across speakergroups (SG's). However, the manner in which such studies arrive at this con…
Identifying and typifying demographic unfairness in phoneme-level embeddings of self-supervised speech recognition models
Felix Herron, Solange Rossato, Alexandre Allauzen +1
Modern automatic speech recognition (ASR) systems have been observed to function better for certain speaker groups (SGs) than others, despite recent gains in overall performance. O…
Where Do Self-Supervised Speech Models Become Unfair?
Felix Herron, Maja Hjuler, Solange Rossato +2
Speech encoder models are known to model members of some speaker groups (SGs) better than others. However, there has been little work in establishing why this occurs on a technolog…
Polynomial Mixing for Efficient Self-supervised Speech Encoders
Eva Feillet, Ryan Whetten, David Picard +1
State-of-the-art speech-to-text models typically employ Transformer-based encoders that model token dependencies via self-attention mechanisms. However, the quadratic complexity of…
Interference Matrix: Quantifying Cross-Lingual Interference in Transformer Encoders
Belen Alastruey, João Maria Janeiro, Alexandre Allauzen +3
In this paper, we present a comprehensive study of language interference in encoder-only Transformer models across 83 languages. We construct an interference matrix by training and…