4 papers
Mapping Written Words to Spoken Words in a Different Language Using Only Visual Grounding
Gabriel Pirlogeanu, Dan Oneata, Horia Cucu +1
In many low-resource settings, even just eliciting speech for data collection is difficult. One promising approach has been to ask speakers to describe images. But how do we build…
Connecting Speech to Words through Images
Gabriel Pirlogeanu, Dan Oneata, Horia Cucu +1
How can we learn the mapping between written words and their spoken counterparts in the absence of explicit textual supervision? We present a visually grounded method for building…
Understanding the strengths and weaknesses of SSL models for audio deepfake model attribution
Gabriel Pîrlogeanu, Adriana Stan, Horia Cucu
Audio deepfake model attribution aims to mitigate the misuse of synthetic speech by identifying the source model responsible for generating a given audio sample, enabling accountab…
Open Source State-Of-the-Art Solution for Romanian Speech Recognition
Gabriel Pirlogeanu, Alexandru-Lucian Georgescu, Horia Cucu
In this work, we present a new state-of-the-art Romanian Automatic Speech Recognition (ASR) system based on NVIDIA's FastConformer architecture--explored here for the first time in…