4 citations · 12 across the 16 of their papers we have counts for
4 papers · 1 filter
Echoes: A semantically-aligned music deepfake detection dataset
Octavian Pascu, Dan Oneata, Horia Cucu +1
We introduce Echoes, a new dataset for music deepfake detection designed for training and benchmarking detectors under realistic and provider-diverse conditions. Echoes comprises 4…
Improving Multimodal Speech Recognition by Data Augmentation and Speech Representations
Dan Oneata, Horia Cucu
Multimodal speech recognition aims to improve the performance of automatic speech recognition (ASR) systems by leveraging additional visual information that is usually associated t…
A Text-to-Speech Pipeline, Evaluation Methodology, and Initial Fine-Tuning Results for Child Speech Synthesis
Rishabh Jain, Mariam Yiwere, Dan Bigioi +2
Speech synthesis has come a long way as current text-to-speech (TTS) models can now generate natural human-sounding speech. However, most of the TTS research focuses on using adult…
Kite: Automatic speech recognition for unmanned aerial vehicles
Dan Oneata, Horia Cucu
This paper addresses the problem of building a speech recognition system attuned to the control of unmanned aerial vehicles (UAVs). Even though UAVs are becoming widespread, the ta…