4 papers · 1 filter
A Toolkit for Joint Speaker Diarization and Identification with Application to Speaker-Attributed ASR
Giovanni Morrone, Enrico Zovato, Fabio Brugnara +2
We present a modular toolkit to perform joint speaker diarization and speaker identification. The toolkit can leverage on multiple models and algorithms which are defined in a conf…
Exploring Spoken Language Identification Strategies for Automatic Transcription of Multilingual Broadcast and Institutional Speech
Martina Valente, Fabio Brugnara, Giovanni Morrone +2
This paper addresses spoken language identification (SLI) and speech recognition of multilingual broadcast and institutional speech, real application scenarios that have been rarel…
Audio-Visual Target Speaker Enhancement on Multi-Talker Environment using Event-Driven Cameras
Ander Arriandiaga, Giovanni Morrone, Luca Pasa +2
We propose a method to address audio-visual target speaker enhancement in multi-talker environments using event-driven cameras. State of the art audio-visual speech separation meth…
An Analysis of Speech Enhancement and Recognition Losses in Limited Resources Multi-talker Single Channel Audio-Visual ASR
Luca Pasa, Giovanni Morrone, Leonardo Badino
In this paper, we analyzed how audio-visual speech enhancement can help to perform the ASR task in a cocktail party scenario. Therefore we considered two simple end-to-end LSTM-bas…