4 papers · 1 filter
ComicsPAP: understanding comic strips by picking the correct panel
Emanuele Vivoli, Artemis Llabrés, Mohamed Ali Souibgui +3
Large multimodal models (LMMs) have made impressive strides in image captioning, VQA, and video comprehension, yet they still struggle with the intricate temporal and spatial cues…
HoloMine: A Synthetic Dataset for Buried Landmines Recognition using Microwave Holographic Imaging
Emanuele Vivoli, Lorenzo Capineri, Marco Bertini
The detection and removal of landmines is a complex and risky task that requires advanced remote sensing techniques to reduce the risk for the professionals involved in this task.…
One missing piece in Vision and Language: A Survey on Comics Understanding
Emanuele Vivoli, Mohamed Ali Souibgui, Andrey Barsky +3
Vision-language models have recently evolved into versatile systems capable of high performance across a range of tasks, such as document understanding, visual question answering,…
CoMix: A Comprehensive Benchmark for Multi-Task Comic Understanding
Emanuele Vivoli, Marco Bertini, Dimosthenis Karatzas
The comic domain is rapidly advancing with the development of single-page analysis and synthesis models. However, evaluation metrics and datasets lag behind, often limited to small…