activity
20202023
most citedAn evaluation of word-level confidence estimation for end-to-end automatic speech recognition

4 citations · 9 across the 6 of their papers we have counts for

collaborators
Showing eess.ASShow all

6 papers · 1 filter

eess.AS20231 cited

Towards generalisable and calibrated synthetic speech detection with self-supervised representations

Octavian Pascu, Adriana Stan, Dan Oneata +2

Generalisation -- the ability of a model to perform well on unseen data -- is crucial for building reliable deepfake detectors. However, recent studies have shown that the current…

eess.AS2023

An analysis on the effects of speaker embedding choice in non auto-regressive TTS

Adriana Stan, Johannah O'Mahony

In this paper we introduce a first attempt on understanding how a non-autoregressive factorised multi-speaker speech synthesis architecture exploits the information present in diff…

eess.AS20233 cited

Residual Information in Deep Speaker Embedding Architectures

Adriana Stan

Speaker embeddings represent a means to extract representative vectorial representations from a speech signal such that the representation pertains to the speaker identity alone. T…

eess.AS20211 cited

Speaker disentanglement in video-to-speech conversion

Dan Oneata, Adriana Stan, Horia Cucu

The task of video-to-speech aims to translate silent video of lip movement to its corresponding audio signal. Previous approaches to this task are generally limited to the case of…

eess.AS20214 cited

An evaluation of word-level confidence estimation for end-to-end automatic speech recognition

Dan Oneata, Alexandru Caranica, Adriana Stan +1

Quantifying the confidence (or conversely the uncertainty) of a prediction is a highly desirable trait of an automatic system, as it improves the robustness and usefulness in downs…

eess.AS2020

RECOApy: Data recording, pre-processing and phonetic transcription for end-to-end speech-based applications

Adriana Stan

Deep learning enables the development of efficient end-to-end speech processing applications while bypassing the need for expert linguistic and signal processing features. Yet, rec…