10 citations · 13 across the 4 of their papers we have counts for
4 papers
A Semi-Automatic Approach to Create Large Gender- and Age-Balanced Speaker Corpora: Usefulness of Speaker Diarization & Identification
Rémi Uro, David Doukhan, Albert Rilliard +4
This paper presents a semi-automatic approach to create a diachronic corpus of voices balanced for speaker's age, gender, and recording period, according to 32 categories (2 gender…
Direct Text to Speech Translation System using Acoustic Units
Victoria Mingote, Pablo Gimeno, Luis Vicente +3
This paper proposes a direct text to speech translation system using discrete acoustic units. This framework employs text in different source languages as input to generate speech…
Joint speech and overlap detection: a benchmark over multiple audio setup and speech domains
Martin Lebourdais, Théo Mariotte, Marie Tahon +5
Voice activity and overlapped speech detection (respectively VAD and OSD) are key pre-processing tasks for speaker diarization. The final segmentation performance highly relies on…
ASR-Generated Text for Language Model Pre-training Applied to Speech Tasks
Valentin Pelloin, Franck Dary, Nicolas Herve +4
We aim at improving spoken language modeling (LM) using very large amount of automatically transcribed speech. We leverage the INA (French National Audiovisual Institute) collectio…