5 papers
Strategic Navigation or Stochastic Search? How Agents and Humans Reason Over Document Collections
Åukasz Borchmann, Jordy Van Landeghem, MichaÅ Turski +12
Multimodal agents offer a promising path to automating complex document-intensive workflows. Yet, a critical question remains: do these agents demonstrate genuine strategic reasoni…
Enhancing Document VQA Models via Retrieval-Augmented Generation
Eric López, Artemis Llabrés, Ernest Valveny
Document Visual Question Answering (Document VQA) must cope with documents that span dozens of pages, yet leading systems still concatenate every page or rely on very large vision-…
ComicsPAP: understanding comic strips by picking the correct panel
Emanuele Vivoli, Artemis Llabrés, Mohamed Ali Souibgui +3
Large multimodal models (LMMs) have made impressive strides in image captioning, VQA, and video comprehension, yet they still struggle with the intricate temporal and spatial cues…
CoSMo: A Multimodal Transformer for Page Stream Segmentation in Comic Books
Marc Serra Ortega, Emanuele Vivoli, Artemis Llabrés +1
This paper introduces CoSMo, a novel multimodal Transformer for Page Stream Segmentation (PSS) in comic books, a critical task for automated content understanding, as it is a neces…
One missing piece in Vision and Language: A Survey on Comics Understanding
Emanuele Vivoli, Mohamed Ali Souibgui, Andrey Barsky +3
Vision-language models have recently evolved into versatile systems capable of high performance across a range of tasks, such as document understanding, visual question answering,…