18 citations · 56 across the 44 of their papers we have counts for
10 papers · 1 filter
Speaker detection in the wild: Lessons learned from JSALT 2019
Paola Garcia, Jesus Villalba, Herve Bredin +21
This paper presents the problems and solutions addressed at the JSALT workshop when using a single microphone for speaker detection in adverse scenarios. The main focus was to tack…
Listen and Fill in the Missing Letters: Non-Autoregressive Transformer for Speech Recognition
Nanxin Chen, Shinji Watanabe, Jesús Villalba +1
Recently very deep transformers have outperformed conventional bi-directional long short-term memory networks by a large margin in speech recognition. However, to put it into produ…
Deep neural networks for emotion recognition combining audio and transcripts
Jaejin Cho, Raghavendra Pappagari, Purva Kulkarni +3
In this paper, we propose to improve emotion recognition by combining acoustic information and conversation transcripts. On the one hand, an LSTM network was used to detect emotion…
Low-Resource Domain Adaptation for Speaker Recognition Using Cycle-GANs
Phani Sankar Nidadavolu, Saurabh Kataria, Jesús Villalba +1
Current speaker recognition technology provides great performance with the x-vector approach. However, performance decreases when the evaluation domain is different from the traini…
Hierarchical Transformers for Long Document Classification
Raghavendra Pappagari, Piotr Żelasko, Jesús Villalba +2
BERT, which stands for Bidirectional Encoder Representations from Transformers, is a recently introduced language representation model based upon the transfer learning paradigm. We…
Unsupervised Feature Enhancement for speaker verification
Phani Sankar Nidadavolu, Saurabh Kataria, Jesús Villalba +2
The task of making speaker verification systems robust to adverse scenarios remain a challenging and an active area of research. We developed an unsupervised feature enhancement ap…