4 papers
NarraBench: A Comprehensive Framework for Narrative Benchmarking
Sil Hamilton, Matthew Wilkens, Andrew Piper
We present NarraBench, a theory-informed taxonomy of narrative-understanding tasks, as well as an associated survey of 78 existing benchmarks in the area. We find significant need…
Too Long, Didn't Model: Decomposing LLM Long-Context Understanding With Novels
Sil Hamilton, Rebecca M. M. Hicke, Matthew Wilkens +1
Although the context length of large language models (LLMs) has increased to millions of tokens, evaluating their effectiveness beyond needle-in-a-haystack approaches has proven di…
Can Language Models Represent the Past without Anachronism?
Ted Underwood, Laura K. Nelson, Matthew Wilkens
Before researchers can use language models to simulate the past, they need to understand the risk of anachronism. We find that prompting a contemporary model with examples of perio…
Tasks and Roles in Legal AI: Data Curation, Annotation, and Verification
Allison Koenecke, Jed Stiglitz, David Mimno +1
The application of AI tools to the legal field feels natural: large legal document collections could be used with specialized AI to improve workflow efficiency for lawyers and amel…