3 papers
cs.CV2025
ComicsPAP: understanding comic strips by picking the correct panel
Emanuele Vivoli, Artemis Llabrés, Mohamed Ali Souibgui +3
Large multimodal models (LMMs) have made impressive strides in image captioning, VQA, and video comprehension, yet they still struggle with the intricate temporal and spatial cues…
cs.CV2025
HoloMine: A Synthetic Dataset for Buried Landmines Recognition using Microwave Holographic Imaging
Emanuele Vivoli, Lorenzo Capineri, Marco Bertini
The detection and removal of landmines is a complex and risky task that requires advanced remote sensing techniques to reduce the risk for the professionals involved in this task.…
cs.CV2025
One missing piece in Vision and Language: A Survey on Comics Understanding
Emanuele Vivoli, Mohamed Ali Souibgui, Andrey Barsky +3
Vision-language models have recently evolved into versatile systems capable of high performance across a range of tasks, such as document understanding, visual question answering,…