182 citations · 212 across the 29 of their papers we have counts for
4 papers · 1 filter
Synthetic Speech Source Tracing using Metric Learning
Dimitrios Koutsianos, Stavros Zacharopoulos, Yannis Panagakis +1
This paper addresses source tracing in synthetic speech-identifying generative systems behind manipulated audio via speaker recognition-inspired pipelines. While prior work focuses…
Investigating Personalization Methods in Text to Music Generation
Manos Plitsis, Theodoros Kouzelis, Georgios Paraskevopoulos +2
In this work, we investigate the personalization of text-to-music diffusion models in a few-shot setting. Motivated by recent advances in the computer vision domain, we are the fir…
Audio-visual video-to-speech synthesis with synthesized input audio
Triantafyllos Kefalas, Yannis Panagakis, Maja Pantic
Video-to-speech synthesis involves reconstructing the speech signal of a speaker from a silent video. The implicit assumption of this task is that the sound signal is either missin…
Large-scale unsupervised audio pre-training for video-to-speech synthesis
Triantafyllos Kefalas, Yannis Panagakis, Maja Pantic
Video-to-speech synthesis is the task of reconstructing the speech signal from a silent video of a speaker. Most established approaches to date involve a two-step process, whereby…