9 citations · 21 across the 12 of their papers we have counts for
3 papers · 1 filter
Leveraging Visual Question Answering to Improve Text-to-Image Synthesis
Stanislav Frolov, Shailza Jolly, Jörn Hees +1
Generating images from textual descriptions has recently attracted a lot of interest. While current models can generate photo-realistic images of individual objects such as birds a…
ESResNet: Environmental Sound Classification Based on Visual Domain Models
Andrey Guzhov, Federico Raue, Jörn Hees +1
Environmental Sound Classification (ESC) is an active research area in the audio domain and has seen a lot of progress in the past years. However, many of the existing approaches a…
P NP, at least in Visual Question Answering
Shailza Jolly, Sebastian Palacio, Joachim Folz +3
In recent years, progress in the Visual Question Answering (VQA) field has largely been driven by public challenges and large datasets. One of the most widely-used of these is the…