3 citations · 3 across the 4 of their papers we have counts for
4 papers
WavLM model ensemble for audio deepfake detection
David Combei, Adriana Stan, Dan Oneata +1
Audio deepfake detection has become a pivotal task over the last couple of years, as many recent speech synthesis and voice cloning systems generate highly realistic speech samples…
An analysis of large speech models-based representations for speech emotion recognition
Adrian Bogdan Stânea, Vlad Striletchi, Cosmin Striletchi +1
Large speech models-derived features have recently shown increased performance over signal-based features across multiple downstream tasks, even when the networks are not finetuned…
An analysis on the effects of speaker embedding choice in non auto-regressive TTS
Adriana Stan, Johannah O'Mahony
In this paper we introduce a first attempt on understanding how a non-autoregressive factorised multi-speaker speech synthesis architecture exploits the information present in diff…
Residual Information in Deep Speaker Embedding Architectures
Adriana Stan
Speaker embeddings represent a means to extract representative vectorial representations from a speech signal such that the representation pertains to the speaker identity alone. T…