18 citations · 63 across the 17 of their papers we have counts for
11 papers · 1 filter
CoCoLex: Confidence-guided Copy-based Decoding for Grounded Legal Text Generation
Santosh T. Y. S. S, Youssef Tarek Elkhayat, Oana Ichim +5
Due to their ability to process long and complex contexts, LLMs can offer key benefits to the Legal domain, but their adoption has been hindered by their tendency to generate unfai…
Where is this coming from? Making groundedness count in the evaluation of Document VQA models
Armineh Nourbakhsh, Siddharth Parekh, Pranav Shetty +3
Document Visual Question Answering (VQA) models have evolved at an impressive rate over the past few years, coming close to or matching human performance on some benchmarks. We arg…
"What is the value of {templates}?" Rethinking Document Information Extraction Datasets for LLMs
Ran Zmigrod, Pranav Shetty, Mathieu Sibue +4
The rise of large language models (LLMs) for visually rich document understanding (VRDU) has kindled a need for prompt-response, document-based datasets. As annotating new datasets…
A Survey on Large Language Models for Critical Societal Domains: Finance, Healthcare, and Law
Zhiyu Zoey Chen, Jing Ma, Xinlu Zhang +7
In the fast-evolving domain of artificial intelligence, large language models (LLMs) such as GPT-3 and GPT-4 are revolutionizing the landscapes of finance, healthcare, and law: dom…
BuDDIE: A Business Document Dataset for Multi-task Information Extraction
Ran Zmigrod, Dongsheng Wang, Mathieu Sibue +10
The field of visually rich document understanding (VRDU) aims to solve a multitude of well-researched NLP tasks in a multi-modal domain. Several datasets exist for research on spec…
TreeForm: End-to-end Annotation and Evaluation for Form Document Parsing
Ran Zmigrod, Zhiqiang Ma, Armineh Nourbakhsh +1
Visually Rich Form Understanding (VRFU) poses a complex research problem due to the documents' highly structured nature and yet highly variable style and content. Current annotatio…