2 citations · 4 across the 11 of their papers we have counts for
3 papers · 1 filter
Improving Explicit Spatial Relationships in Text-to-Image Generation through an Automatically Derived Dataset
Ander Salaberria, Gorka Azkune, Oier Lopez de Lacalle +3
Existing work has observed that current text-to-image systems do not accurately reflect explicit spatial relations between objects such as 'left of' or 'below'. We hypothesize that…
Learning Action Changes by Measuring Verb-Adverb Textual Relationships
Davide Moltisanti, Frank Keller, Hakan Bilen +1
The goal of this work is to understand the way actions are performed in videos. That is, given a video, we aim to predict an adverb indicating a modification applied to the action…
Learn2Augment: Learning to Composite Videos for Data Augmentation in Action Recognition
Shreyank N Gowda, Marcus Rohrbach, Frank Keller +1
We address the problem of data augmentation for video action recognition. Standard augmentation strategies in video are hand-designed and sample the space of possible augmented dat…