Publications (4)
AlloVera: A Multilingual Allophone Database
David R. Mortensen, Xinjian Li, Patrick Littell +6
We introduce a new resource, AlloVera, which provides mappings from 218 allophones to phonemes for 14 languages. Phonemes are contrastive phonological units, and allophones are the…
User-friendly automatic transcription of low-resource languages: Plugging ESPnet into Elpis
Oliver Adams, Benjamin Galliot, Guillaume Wisniewski +11
This paper reports on progress integrating the speech recognition toolkit ESPnet into Elpis, a web front-end originally designed to provide access to the Kaldi automatic speech rec…
Establishing degrees of closeness between audio recordings along different dimensions using large-scale cross-lingual models
Maxime Fily, Guillaume Wisniewski, Severine Guillaume +2
In the highly constrained context of low-resource language studies, we explore vector representations of speech from a pretrained model to determine their level of abstraction with…
From `Snippet-lects' to Doculects and Dialects: Leveraging Neural Representations of Speech for Placing Audio Signals in a Language Landscape
Séverine Guillaume, Guillaume Wisniewski, Alexis Michaud
XLSR-53 a multilingual model of speech, builds a vector representation from audio, which allows for a range of computational treatments. The experiments reported here use this neur…