papers

Publications (14)

cs.CL2024

Core: Robust Factual Precision with Informative Sub-Claim Identification

Zhengping Jiang, Jingyu Zhang, Nathaniel Weir +6

Hallucinations pose a challenge to the application of large language models (LLMs) thereby motivating the development of metrics to evaluate factual precision. We observe that popu…

cs.CL2021

Gradual Fine-Tuning for Low-Resource Domain Adaptation

Haoran Xu, Seth Ebner, Mahsa Yarmohammadi +3

Fine-tuning is known to improve NLP models by adapting an initial model trained on more plentiful but less domain-salient examples to data in a target domain. Such domain adaptatio…

cs.CL2026

The Effect of Scripts and Formats on LLM Numeracy

Varshini Reddy, Craig W. Schmidt, Seth Ebner +3

Large language models (LLMs) have achieved impressive proficiency in basic arithmetic, rivaling human-level performance on standard numerical tasks. However, little attention has b…

cs.CL2026

No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding

Michael Krumdick, Charles Lovering, Varshini Reddy +2

Reliable evaluation of large language models (LLMs) is critical as their deployment rapidly expands, particularly in high-stakes domains such as business and finance. The LLM-as-a-…

cs.CL2020

Multi-Sentence Argument Linking

Seth Ebner, Patrick Xia, Ryan Culkin +2

We present a novel document-level model for finding argument spans that fill an event's roles, connecting related ideas in sentence-level semantic role labeling and coreference res…

cs.CL2022

An Augmentation Strategy for Visually Rich Documents

Jing Xie, James B. Wendt, Yichao Zhou +2

Many business workflows require extracting important fields from form-like documents (e.g. bank statements, bills of lading, purchase orders, etc.). Recent techniques for automatin…

cs.CL2024

A Closer Look at Claim Decomposition

Miriam Wanner, Seth Ebner, Zhengping Jiang +2

As generated text becomes more commonplace, it is increasingly important to evaluate how well-supported such text is by external knowledge sources. Many approaches for evaluating t…

cs.CL2025

On Finding Inconsistencies in Documents

Charles J. Lovering, Seth Ebner, Brandon Smock +5

Professionals in academia, law, and finance audit their documents because inconsistencies can result in monetary, reputational, and scientific costs. Language models (LMs) have the…

cs.CL2026

Tokenization with Split Trees

Craig W. Schmidt, Michael Krumdick, Adam Wiemerslage +4

We introduce Tokenization with Split Trees (ToaST), a subword tokenization method that directly optimizes compression under a new recursive inference procedure. ToaST greedily spli…

cs.CL2026

Cost-Efficient Estimation of General Abilities Across Benchmarks

Michael Krumdick, Adam Wiemerslage, Seth Ebner +2

Thousands of diverse benchmarks have been developed to measure the quality of large language models (LLMs). Yet prior work has demonstrated that LLM performance is often sufficient…

cs.SI2019

An Exact No Free Lunch Theorem for Community Detection

Arya D. McCarthy, Tongfei Chen, Seth Ebner

A precondition for a No Free Lunch theorem is evaluation with a loss function which does not assume a priori superiority of some outputs over others. A previous result for communit…

cs.CL2021

Everything Is All It Takes: A Multipronged Strategy for Zero-Shot Cross-Lingual Information Extraction

Mahsa Yarmohammadi, Shijie Wu, Marc Marone +10

Zero-shot cross-lingual information extraction (IE) describes the construction of an IE model for some target language, given existing annotations exclusively in some other languag…

cs.CL2020

Reading the Manual: Event Extraction as Definition Comprehension

Yunmo Chen, Tongfei Chen, Seth Ebner +2

We ask whether text understanding has progressed to where we may extract event information through incremental refinement of bleached statements derived from annotation manuals. Su…

cs.AI2025

Language Model Probabilities are Not Calibrated in Numeric Contexts

Charles Lovering, Michael Krumdick, Viet Dac Lai +5

Some statements have one well-defined continuation (e.g., "the Eiffel Tower is in [Paris]"), whereas others have a natural distribution over multiple options (e.g., "the weighted c…