4 citations · 6 across the 4 of their papers we have counts for
6 papers
Speech inpainting: Context-based speech synthesis guided by video
Juan F. Montesinos, Daniel Michelsanti, Gloria Haro +2
Audio and visual modalities are inherently connected in speech signals: lip movements and facial expressions are correlated with speech sounds. This motivates studies that incorpor…
VocaLiST: An Audio-Visual Synchronisation Model for Lips and Voices
Venkatesh S. Kadandale, Juan F. Montesinos, Gloria Haro
In this paper, we address the problem of lip-voice synchronisation in videos containing human face and voice. Our approach is based on determining if the lips motion and the voice…
VoViT: Low Latency Graph-based Audio-Visual Voice Separation Transformer
Juan F. Montesinos, Venkatesh S. Kadandale, Gloria Haro
This paper presents an audio-visual approach for voice separation which produces state-of-the-art results at a low latency in two scenarios: speech and singing voice. The model is…
A cappella: Audio-visual Singing Voice Separation
Juan F. Montesinos, Venkatesh S. Kadandale, Gloria Haro
The task of isolating a target singing voice in music videos has useful applications. In this work, we explore the single-channel singing voice separation problem from a multimodal…
Solos: A Dataset for Audio-Visual Music Analysis
Juan F. Montesinos, Olga Slizovskaia, Gloria Haro
In this paper, we present a new dataset of music performance videos which can be used for training machine learning methods for multiple tasks such as audio-visual blind source sep…
Multi-channel U-Net for Music Source Separation
Venkatesh S. Kadandale, Juan F. Montesinos, Gloria Haro +1
A fairly straightforward approach for music source separation is to train independent models, wherein each model is dedicated for estimating only a specific source. Training a sing…