1 citations · 1 across the 4 of their papers we have counts for
4 papers
RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors
Liam Dugan, Alyssa Hwang, Filip Trhlik +5
Many commercial and open-source models claim to detect machine-generated text with extremely high accuracy (99% or more). However, very few of these detectors are evaluated on shar…
Large Language Models Fall Short: Understanding Complex Relationships in Detective Narratives
Runcong Zhao, Qinglin Zhu, Hainiu Xu +4
Existing datasets for narrative understanding often fail to represent the complexity and uncertainty of relationships in real-life social scenarios. To address this gap, we introdu…
Exploring the Curious Case of Code Prompts
Li Zhang, Liam Dugan, Hainiu Xu +1
Recent work has shown that prompting language models with code-like representations of natural language leads to performance improvements on structured reasoning tasks. However, su…
Causal Reasoning of Entities and Events in Procedural Texts
Li Zhang, Hainiu Xu, Yue Yang +4
Entities and events are crucial to natural language reasoning and common in procedural texts. Existing work has focused either exclusively on entity state tracking (e.g., whether a…