Publications (14)
Core: Robust Factual Precision with Informative Sub-Claim Identification
Zhengping Jiang, Jingyu Zhang, Nathaniel Weir +6
Hallucinations pose a challenge to the application of large language models (LLMs) thereby motivating the development of metrics to evaluate factual precision. We observe that popu…
Gradual Fine-Tuning for Low-Resource Domain Adaptation
Haoran Xu, Seth Ebner, Mahsa Yarmohammadi +3
Fine-tuning is known to improve NLP models by adapting an initial model trained on more plentiful but less domain-salient examples to data in a target domain. Such domain adaptatio…
The Effect of Scripts and Formats on LLM Numeracy
Varshini Reddy, Craig W. Schmidt, Seth Ebner +3
Large language models (LLMs) have achieved impressive proficiency in basic arithmetic, rivaling human-level performance on standard numerical tasks. However, little attention has b…
No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding
Michael Krumdick, Charles Lovering, Varshini Reddy +2
Reliable evaluation of large language models (LLMs) is critical as their deployment rapidly expands, particularly in high-stakes domains such as business and finance. The LLM-as-a-…
Multi-Sentence Argument Linking
Seth Ebner, Patrick Xia, Ryan Culkin +2
We present a novel document-level model for finding argument spans that fill an event's roles, connecting related ideas in sentence-level semantic role labeling and coreference res…
An Augmentation Strategy for Visually Rich Documents
Jing Xie, James B. Wendt, Yichao Zhou +2
Many business workflows require extracting important fields from form-like documents (e.g. bank statements, bills of lading, purchase orders, etc.). Recent techniques for automatin…
A Closer Look at Claim Decomposition
Miriam Wanner, Seth Ebner, Zhengping Jiang +2
As generated text becomes more commonplace, it is increasingly important to evaluate how well-supported such text is by external knowledge sources. Many approaches for evaluating t…
On Finding Inconsistencies in Documents
Charles J. Lovering, Seth Ebner, Brandon Smock +5
Professionals in academia, law, and finance audit their documents because inconsistencies can result in monetary, reputational, and scientific costs. Language models (LMs) have the…
Tokenization with Split Trees
Craig W. Schmidt, Michael Krumdick, Adam Wiemerslage +4
We introduce Tokenization with Split Trees (ToaST), a subword tokenization method that directly optimizes compression under a new recursive inference procedure. ToaST greedily spli…
Cost-Efficient Estimation of General Abilities Across Benchmarks
Michael Krumdick, Adam Wiemerslage, Seth Ebner +2
Thousands of diverse benchmarks have been developed to measure the quality of large language models (LLMs). Yet prior work has demonstrated that LLM performance is often sufficient…
An Exact No Free Lunch Theorem for Community Detection
Arya D. McCarthy, Tongfei Chen, Seth Ebner
A precondition for a No Free Lunch theorem is evaluation with a loss function which does not assume a priori superiority of some outputs over others. A previous result for communit…
Everything Is All It Takes: A Multipronged Strategy for Zero-Shot Cross-Lingual Information Extraction
Mahsa Yarmohammadi, Shijie Wu, Marc Marone +10
Zero-shot cross-lingual information extraction (IE) describes the construction of an IE model for some target language, given existing annotations exclusively in some other languag…
Reading the Manual: Event Extraction as Definition Comprehension
Yunmo Chen, Tongfei Chen, Seth Ebner +2
We ask whether text understanding has progressed to where we may extract event information through incremental refinement of bleached statements derived from annotation manuals. Su…
Language Model Probabilities are Not Calibrated in Numeric Contexts
Charles Lovering, Michael Krumdick, Viet Dac Lai +5
Some statements have one well-defined continuation (e.g., "the Eiffel Tower is in [Paris]"), whereas others have a natural distribution over multiple options (e.g., "the weighted c…