BoundingDocs: a Unified Dataset for Document Question Answering with Spatial Annotations
arXiv:2501.03403 · doi:10.1007/s10032-025-00563-5
Abstract
We present a unified dataset for document Question-Answering (QA), which is obtained combining several public datasets related to Document AI and visually rich document understanding (VRDU). Our main contribution is twofold: on the one hand we reformulate existing Document AI tasks, such as Information Extraction (IE), into a Question-Answering task, making it a suitable resource for training and evaluating Large Language Models; on the other hand, we release the OCR of all the documents and include the exact position of the answer to be found in the document image as a bounding box. Using this dataset, we explore the impact of different prompting techniques (that might include bounding box information) on the performance of open-weight models, identifying the most effective approaches for document comprehension.
References in corpus (12)
- LLaMA: Open and Efficient Foundation Language Models
- ICDAR2019 Competition on Scanned Receipt OCR and Information Extraction
- Mistral 7B
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- VRDU: A Benchmark for Visually-rich Document Understanding
- Layout and Task Aware Instruction Prompt for Zero-shot Document Image Question Answering
- "What is the value of {templates}?" Rethinking Document Information Extraction Datasets for LLMs
- ANLS* -- A Universal Document Processing Metric for Generative Large Language Models
- mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding
- BigDocs: An Open Dataset for Training Multimodal Models on Document and Code Tasks
- FATURA: A Multi-Layout Invoice Image Dataset for Document Analysis and Understanding