3 papers
cs.CL2025
BoundingDocs: a Unified Dataset for Document Question Answering with Spatial Annotations
Simone Giovannini, Fabio Coppini, Andrea Gemelli +1
We present a unified dataset for document Question-Answering (QA), which is obtained combining several public datasets related to Document AI and visually rich document understandi…
cs.CL2025
Hierarchical structure understanding in complex tables with VLLMs: a benchmark and experiments
Luca Bindini, Simone Giovannini, Simone Marinai +2
This work investigates the ability of Vision Large Language Models (VLLMs) to understand and interpret the structure of tables in scientific articles. Specifically, we explore whet…
cs.CL2025
Towards Reliable and Interpretable Document Question Answering via VLMs
Alessio Chen, Simone Giovannini, Andrea Gemelli +2
Vision-Language Models (VLMs) have shown strong capabilities in document understanding, particularly in identifying and extracting textual information from complex documents. Despi…