4 citations · 7 across the 6 of their papers we have counts for
6 papers
Improving Multimodal Speech Recognition by Data Augmentation and Speech Representations
Dan Oneata, Horia Cucu
Multimodal speech recognition aims to improve the performance of automatic speech recognition (ASR) systems by leveraging additional visual information that is usually associated t…
A Text-to-Speech Pipeline, Evaluation Methodology, and Initial Fine-Tuning Results for Child Speech Synthesis
Rishabh Jain, Mariam Yiwere, Dan Bigioi +2
Speech synthesis has come a long way as current text-to-speech (TTS) models can now generate natural human-sounding speech. However, most of the TTS research focuses on using adult…
Speaker disentanglement in video-to-speech conversion
Dan Oneata, Adriana Stan, Horia Cucu
The task of video-to-speech aims to translate silent video of lip movement to its corresponding audio signal. Previous approaches to this task are generally limited to the case of…
An evaluation of word-level confidence estimation for end-to-end automatic speech recognition
Dan Oneata, Alexandru Caranica, Adriana Stan +1
Quantifying the confidence (or conversely the uncertainty) of a prediction is a highly desirable trait of an automatic system, as it improves the robustness and usefulness in downs…
The Quo Vadis submission at Traffic4cast 2019
Dan Oneata, Cosmin George Alexandru, Marius Stanescu +4
We describe the submission of the Quo Vadis team to the Traffic4cast competition, which was organized as part of the NeurIPS 2019 series of challenges. Our system consists of a tem…
Kite: Automatic speech recognition for unmanned aerial vehicles
Dan Oneata, Horia Cucu
This paper addresses the problem of building a speech recognition system attuned to the control of unmanned aerial vehicles (UAVs). Even though UAVs are becoming widespread, the ta…