6 papers · 1 filter
Weird Generalization is Weirdly Brittle
Miriam Wanner, Hannah Collison, William Jurayj +3
Weird generalization is a phenomenon in which models fine-tuned on data from a narrow domain (e.g. insecure code) develop surprising traits that manifest even outside that domain (…
All Claims Are Equal, but Some Claims Are More Equal Than Others: Importance-Sensitive Factuality Evaluation of LLM Generations
Miriam Wanner, Leif Azzopardi, Paul Thomas +3
Existing methods for evaluating the factuality of large language model (LLM) responses treat all claims as equally important. This results in misleading evaluations when vital info…
How Grounded is Wikipedia? A Study on Structured Evidential Support and Retrieval
William Walden, Kathryn Ricci, Miriam Wanner +4
Wikipedia is a critical resource for modern NLP, serving as a rich repository of up-to-date and citation-backed information on a wide variety of subjects. The reliability of Wikipe…
CLAIMCHECK: How Grounded are LLM Critiques of Scientific Papers?
Jiefu Ou, William Gantt Walden, Kate Sanders +13
A core part of scientific peer review involves providing expert critiques that directly assess the scientific claims a paper makes. While it is now possible to automatically genera…
DnDScore: Decontextualization and Decomposition for Factuality Verification in Long-Form Text Generation
Miriam Wanner, Benjamin Van Durme, Mark Dredze
The decompose-then-verify strategy for verification of Large Language Model (LLM) generations decomposes claims that are then independently verified. Decontextualization augments t…
Revisiting the Effects of Leakage on Dependency Parsing
Nathaniel Krasner, Miriam Wanner, Antonios Anastasopoulos
Recent work by Søgaard (2020) showed that, treebank size aside, overlap between training and test graphs (termed leakage) explains more of the observed variation in dependency pars…