24 citations · 24 across the 1 of their papers we have counts for
1 paper
William Havard, Laurent Besacier, Olivier Rosec
This paper presents an augmentation of MSCOCO dataset where speech is added to image and text. Speech captions are generated using text-to-speech (TTS) synthesis resulting in 616,7…