Showing cs.CVShow all
3 papers · 1 filter
cs.CV2024
CIC-BART-SSA: Controllable Image Captioning with Structured Semantic Augmentation
Kalliopi Basioti, Mohamed A. Abdelsalam, Federico Fancellu +2
Controllable Image Captioning (CIC) aims at generating natural language descriptions for an image, conditioned on information provided by end users, e.g., regions, entities or even…
cs.CV2023
GePSAn: Generative Procedure Step Anticipation in Cooking Videos
Mohamed Ashraf Abdelsalam, Samrudhdhi B. Rangrej, Isma Hadji +3
We study the problem of future step anticipation in procedural videos. Given a video of an ongoing procedural activity, we predict a plausible next procedure step described in rich…
cs.CV2022
Visual Semantic Parsing: From Images to Abstract Meaning Representation
Mohamed Ashraf Abdelsalam, Zhan Shi, Federico Fancellu +4
The success of scene graphs for visual scene understanding has brought attention to the benefits of abstracting a visual input (e.g., image) into a structured representation, where…