3 papers
eess.AS2019
Speaker detection in the wild: Lessons learned from JSALT 2019
Paola Garcia, Jesus Villalba, Herve Bredin +21
This paper presents the problems and solutions addressed at the JSALT workshop when using a single microphone for speaker detection in adverse scenarios. The main focus was to tack…
eess.AS2019
pyannote.audio: neural building blocks for speaker diarization
Hervé Bredin, Ruiqing Yin, Juan Manuel Coria +7
We introduce pyannote.audio, an open-source toolkit written in Python for speaker diarization. Based on PyTorch machine learning framework, it provides a set of trainable end-to-en…
eess.AS2019
End-to-end Domain-Adversarial Voice Activity Detection
Marvin Lavechin, Marie-Philippe Gill, Ruben Bousbib +2
Voice activity detection is the task of detecting speech regions in a given audio stream or recording. First, we design a neural network combining trainable filters and recurrent l…