2 citations · 2 across the 3 of their papers we have counts for
6 papers · 1 filter
A Toolkit for Joint Speaker Diarization and Identification with Application to Speaker-Attributed ASR
Giovanni Morrone, Enrico Zovato, Fabio Brugnara +2
We present a modular toolkit to perform joint speaker diarization and speaker identification. The toolkit can leverage on multiple models and algorithms which are defined in a conf…
Exploring Spoken Language Identification Strategies for Automatic Transcription of Multilingual Broadcast and Institutional Speech
Martina Valente, Fabio Brugnara, Giovanni Morrone +2
This paper addresses spoken language identification (SLI) and speech recognition of multilingual broadcast and institutional speech, real application scenarios that have been rarel…
Conversational Speech Separation: an Evaluation Study for Streaming Applications
Giovanni Morrone, Samuele Cornell, Enrico Zovato +2
Continuous speech separation (CSS) is a recently proposed framework which aims at separating each speaker from an input mixture signal in a streaming fashion. Hereafter we perform…
Audio-Visual Speech Inpainting with Deep Learning
Giovanni Morrone, Daniel Michelsanti, Zheng-Hua Tan +1
In this paper, we present a deep-learning-based framework for audio-visual speech inpainting, i.e., the task of restoring the missing parts of an acoustic speech signal from reliab…
Audio-Visual Target Speaker Enhancement on Multi-Talker Environment using Event-Driven Cameras
Ander Arriandiaga, Giovanni Morrone, Luca Pasa +2
We propose a method to address audio-visual target speaker enhancement in multi-talker environments using event-driven cameras. State of the art audio-visual speech separation meth…
An Analysis of Speech Enhancement and Recognition Losses in Limited Resources Multi-talker Single Channel Audio-Visual ASR
Luca Pasa, Giovanni Morrone, Leonardo Badino
In this paper, we analyzed how audio-visual speech enhancement can help to perform the ASR task in a cocktail party scenario. Therefore we considered two simple end-to-end LSTM-bas…