5 citations · 6 across the 2 of their papers we have counts for
3 papers
cs.CV2022★ 5 cited
Contrastive Language-Action Pre-training for Temporal Localization
Mengmeng Xu, Erhan Gundogdu, Maksim Lapin +3
Long-form video understanding requires designing approaches that are able to temporally localize activities or language. End-to-end training for such tasks is limited by the comput…
cs.CV2021★ 1 cited
Revamping Cross-Modal Recipe Retrieval with Hierarchical Transformers and Self-supervised Learning
Amaia Salvador, Erhan Gundogdu, Loris Bazzani +1
Cross-modal recipe retrieval has recently gained substantial attention due to the importance of food in people's lives, as well as the availability of vast amounts of digital cooki…
cs.CV2018
Image Captioning as Neural Machine Translation Task in SOCKEYE
Loris Bazzani, Tobias Domhan, Felix Hieber
Image captioning is an interdisciplinary research problem that stands between computer vision and natural language processing. The task is to generate a textual description of the…