activity
20242026
most citedComicsPAP: understanding comic strips by picking the correct panel

2 citations · 2 across the 7 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2025

Enhancing Document VQA Models via Retrieval-Augmented Generation

Eric López, Artemis Llabrés, Ernest Valveny

Document Visual Question Answering (Document VQA) must cope with documents that span dozens of pages, yet leading systems still concatenate every page or rely on very large vision-…

cs.CV2025

CoSMo: A Multimodal Transformer for Page Stream Segmentation in Comic Books

Marc Serra Ortega, Emanuele Vivoli, Artemis Llabrés +1

This paper introduces CoSMo, a novel multimodal Transformer for Page Stream Segmentation (PSS) in comic books, a critical task for automated content understanding, as it is a neces…

cs.CV2025★ 2 cited

ComicsPAP: understanding comic strips by picking the correct panel

Emanuele Vivoli, Artemis Llabrés, Mohamed Ali Souibgui +3

Large multimodal models (LMMs) have made impressive strides in image captioning, VQA, and video comprehension, yet they still struggle with the intricate temporal and spatial cues…

cs.CV2024

One missing piece in Vision and Language: A Survey on Comics Understanding

Emanuele Vivoli, Mohamed Ali Souibgui, Andrey Barsky +3

Vision-language models have recently evolved into versatile systems capable of high performance across a range of tasks, such as document understanding, visual question answering,…

cs.CV2024

Image-text matching for large-scale book collections

Artemis Llabrés, Arka Ujjal Dey, Dimosthenis Karatzas +1

We address the problem of detecting and mapping all books in a collection of images to entries in a given book catalogue. Instead of performing independent retrieval for each book…