14 papers
Don't `Well, Actually' Me Unless You Know What You're Talking About: Weak Presupposition Verification Degrades General QA Performance
Shenran Wang, Vered Shwartz, Hila Gonen
False-presupposition QA (FPQA) tests LLMs on their ability to identify false presuppositions in questions and abstain or correct them rather than reinforcing false assumptions. The…
LLMs Infer Cultural Context but Fail to Apply It When Responding
Yisong Miao, Jian Zhu, Vered Shwartz
Recent work has shown that LLMs overrepresent dominant cultures, particularly Western ones, while marginalizing others. We investigate whether this affects models' ability to gener…
The CIFAR Synthetic Evidence Corpus for Detecting AI-Generated Evidence
Kelly McConvey, Jalehsadat Mahdavimoghaddam, Nima Jamali +8
The growing ability of generative models to produce realistic documents poses a direct challenge to evidentiary workflows in the justice system and the courts, where decisions incr…
Quantifying Media Representation Dynamics Across 25 Years of News Reporting on Policing-related Deaths
Farhan Samir, Jappun Dhillon, Meghna Ravikumar +2
We perform the largest known computational analysis of Canadian news narratives about police-involved deaths, spanning 4,000 articles from the last quarter-century. We develop a no…
CanLegalRAGBench: Evaluating Retrieval-Augmented Generation on Canadian Case Law
Ethan Zhao, Maksym Taranukhin, Wei Cui +2
RAG-based legal assistants have been growing in popularity, but LLM hallucinations remain a key issue and potentially undermines justice. While benchmarks have been developed to ev…
InfoGatherer: Principled Information Seeking via Evidence Retrieval and Strategic Questioning
Maksym Taranukhin, Shuyue Stella Li, Evangelos Milios +3
LLMs are increasingly deployed in high-stakes domains such as medical triage and legal assistance, often as document-grounded QA systems in which a user provides a description, rel…