5 papers
ComicsPAP: understanding comic strips by picking the correct panel
Emanuele Vivoli, Artemis Llabrés, Mohamed Ali Souibgui +3
Large multimodal models (LMMs) have made impressive strides in image captioning, VQA, and video comprehension, yet they still struggle with the intricate temporal and spatial cues…
CoSMo: A Multimodal Transformer for Page Stream Segmentation in Comic Books
Marc Serra Ortega, Emanuele Vivoli, Artemis Llabrés +1
This paper introduces CoSMo, a novel multimodal Transformer for Page Stream Segmentation (PSS) in comic books, a critical task for automated content understanding, as it is a neces…
HoloMine: A Synthetic Dataset for Buried Landmines Recognition using Microwave Holographic Imaging
Emanuele Vivoli, Lorenzo Capineri, Marco Bertini
The detection and removal of landmines is a complex and risky task that requires advanced remote sensing techniques to reduce the risk for the professionals involved in this task.…
One missing piece in Vision and Language: A Survey on Comics Understanding
Emanuele Vivoli, Mohamed Ali Souibgui, Andrey Barsky +3
Vision-language models have recently evolved into versatile systems capable of high performance across a range of tasks, such as document understanding, visual question answering,…
CoMix: A Comprehensive Benchmark for Multi-Task Comic Understanding
Emanuele Vivoli, Marco Bertini, Dimosthenis Karatzas
The comic domain is rapidly advancing with the development of single-page analysis and synthesis models. However, evaluation metrics and datasets lag behind, often limited to small…