5 citations · 9 across the 3 of their papers we have counts for
3 papers
iEdit: Localised Text-guided Image Editing with Weak Supervision
Rumeysa Bodur, Erhan Gundogdu, Binod Bhattarai +3
Diffusion models (DMs) can generate realistic images with text guidance using large-scale datasets. However, they demonstrate limited controllability in the output space of the gen…
Contrastive Language-Action Pre-training for Temporal Localization
Mengmeng Xu, Erhan Gundogdu, Maksim Lapin +3
Long-form video understanding requires designing approaches that are able to temporally localize activities or language. End-to-end training for such tasks is limited by the comput…
Revamping Cross-Modal Recipe Retrieval with Hierarchical Transformers and Self-supervised Learning
Amaia Salvador, Erhan Gundogdu, Loris Bazzani +1
Cross-modal recipe retrieval has recently gained substantial attention due to the importance of food in people's lives, as well as the availability of vast amounts of digital cooki…