4 citations · 9 across the 6 of their papers we have counts for
6 papers · 1 filter
Towards generalisable and calibrated synthetic speech detection with self-supervised representations
Octavian Pascu, Adriana Stan, Dan Oneata +2
Generalisation -- the ability of a model to perform well on unseen data -- is crucial for building reliable deepfake detectors. However, recent studies have shown that the current…
An analysis on the effects of speaker embedding choice in non auto-regressive TTS
Adriana Stan, Johannah O'Mahony
In this paper we introduce a first attempt on understanding how a non-autoregressive factorised multi-speaker speech synthesis architecture exploits the information present in diff…
Residual Information in Deep Speaker Embedding Architectures
Adriana Stan
Speaker embeddings represent a means to extract representative vectorial representations from a speech signal such that the representation pertains to the speaker identity alone. T…
Speaker disentanglement in video-to-speech conversion
Dan Oneata, Adriana Stan, Horia Cucu
The task of video-to-speech aims to translate silent video of lip movement to its corresponding audio signal. Previous approaches to this task are generally limited to the case of…
An evaluation of word-level confidence estimation for end-to-end automatic speech recognition
Dan Oneata, Alexandru Caranica, Adriana Stan +1
Quantifying the confidence (or conversely the uncertainty) of a prediction is a highly desirable trait of an automatic system, as it improves the robustness and usefulness in downs…
RECOApy: Data recording, pre-processing and phonetic transcription for end-to-end speech-based applications
Adriana Stan
Deep learning enables the development of efficient end-to-end speech processing applications while bypassing the need for expert linguistic and signal processing features. Yet, rec…