activity
20242026
collaborators

5 papers

cs.CL2026

ExStrucTiny: A Benchmark for Schema-Variable Structured Information Extraction from Document Images

Mathieu Sibue, Andres Muñoz Garza, Samuel Mensah +4

Enterprise documents, such as forms and reports, embed critical information for downstream applications like data archiving, automated workflows, and analytics. Although generalist…

cs.CL2025

Perturb Your Data: Paraphrase-Guided Training Data Watermarking

Pranav Shetty, Mirazul Haque, Petr Babkin +3

Training data detection is critical for enforcing copyright and data licensing, as Large Language Models (LLM) are trained on massive text corpora scraped from the internet. We pre…

cs.CL2025

CoCoLex: Confidence-guided Copy-based Decoding for Grounded Legal Text Generation

Santosh T. Y. S. S, Youssef Tarek Elkhayat, Oana Ichim +5

Due to their ability to process long and complex contexts, LLMs can offer key benefits to the Legal domain, but their adoption has been hindered by their tendency to generate unfai…

cs.CL2025

Where is this coming from? Making groundedness count in the evaluation of Document VQA models

Armineh Nourbakhsh, Siddharth Parekh, Pranav Shetty +3

Document Visual Question Answering (VQA) models have evolved at an impressive rate over the past few years, coming close to or matching human performance on some benchmarks. We arg…

cs.CL2024

"What is the value of {templates}?" Rethinking Document Information Extraction Datasets for LLMs

Ran Zmigrod, Pranav Shetty, Mathieu Sibue +4

The rise of large language models (LLMs) for visually rich document understanding (VRDU) has kindled a need for prompt-response, document-based datasets. As annotating new datasets…