9 citations · 9 across the 3 of their papers we have counts for
3 papers
cs.CL2025
The Oracle Has Spoken: A Multi-Aspect Evaluation of Dialogue in Pythia
Zixun Chen, Petr Babkin, Akshat Gupta +2
Dialogue is one of the landmark abilities of large language models (LLMs). Despite its ubiquity, few studies actually distinguish specific ingredients underpinning dialogue behavio…
cs.CL2024
BuDDIE: A Business Document Dataset for Multi-task Information Extraction
Ran Zmigrod, Dongsheng Wang, Mathieu Sibue +10
The field of visually rich document understanding (VRDU) aims to solve a multitude of well-researched NLP tasks in a multi-modal domain. Several datasets exist for research on spec…
cs.CL2023★ 9 cited
DocLLM: A layout-aware generative language model for multimodal document understanding
Dongsheng Wang, Natraj Raman, Mathieu Sibue +6
Enterprise documents such as forms, invoices, receipts, reports, contracts, and other similar records, often carry rich semantics at the intersection of textual and spatial modalit…