5 papers
Expect the Unexpected? Testing the Surprisal of Salient Entities
Jessica Lin, Amir Zeldes
Previous work examining the Uniform Information Density (UID) hypothesis has shown that while information as measured by surprisal metrics is distributed more or less evenly across…
DiscoTrack: A Multilingual LLM Benchmark for Discourse Tracking
Lanni Bu, Lauren Levine, Amir Zeldes
Recent LLM benchmarks have tested models on a range of phenomena, but are still focused primarily on natural language understanding for extraction of explicit information, such as…
What makes an entity salient in discourse?
Amir Zeldes, Jessica Lin
Entities in discourse vary in salience: main participants, objects and locations stay prominent, while others are quickly forgotten, raising questions about how humans signal and i…
GUM-SAGE: A Novel Dataset and Approach for Graded Entity Salience Prediction
Jessica Lin, Amir Zeldes
Determining and ranking the most salient entities in a text is critical for user-facing systems, especially as users increasingly rely on models to interpret long documents they on…
GDTB: Genre Diverse Data for English Shallow Discourse Parsing across Modalities, Text Types, and Domains
Yang Janet Liu, Tatsuya Aoyama, Wesley Scivetti +6
Work on shallow discourse parsing in English has focused on the Wall Street Journal corpus, the only large-scale dataset for the language in the PDTB framework. However, the data i…