8 papers
Can Reasoning Models Detect Changes to their Chains of Thought?
Sathvik Napa, Utkarsh Singh, Chengyuan Xue +2
There are many reasons one may want to edit a model's chain of thought (CoT) -- e.g., to prefill it with reasoning from a stronger model or to remove steps that may yield unsafe ou…
Weird Generalization is Weirdly Brittle
Miriam Wanner, Hannah Collison, William Jurayj +3
Weird generalization is a phenomenon in which models fine-tuned on data from a narrow domain (e.g. insecure code) develop surprising traits that manifest even outside that domain (…
Reasoning Models Will Sometimes Lie About Their Reasoning
William Walden, Miriam Wanner
Hint-based faithfulness evaluations have established that Large Reasoning Models (LRMs) may not say what they think: they do not always volunteer information about how key parts of…
All Claims Are Equal, but Some Claims Are More Equal Than Others: Importance-Sensitive Factuality Evaluation of LLM Generations
Miriam Wanner, Leif Azzopardi, Paul Thomas +3
Existing methods for evaluating the factuality of large language model (LLM) responses treat all claims as equally important. This results in misleading evaluations when vital info…
How Grounded is Wikipedia? A Study on Structured Evidential Support and Retrieval
William Walden, Kathryn Ricci, Miriam Wanner +4
Wikipedia is a critical resource for modern NLP, serving as a rich repository of up-to-date and citation-backed information on a wide variety of subjects. The reliability of Wikipe…
CLAIMCHECK: How Grounded are LLM Critiques of Scientific Papers?
Jiefu Ou, William Gantt Walden, Kate Sanders +13
A core part of scientific peer review involves providing expert critiques that directly assess the scientific claims a paper makes. While it is now possible to automatically genera…