3 papers
cs.AI2026
Can Reasoning Models Detect Changes to their Chains of Thought?
Sathvik Napa, Utkarsh Singh, Chengyuan Xue +2
There are many reasons one may want to edit a model's chain of thought (CoT) -- e.g., to prefill it with reasoning from a stronger model or to remove steps that may yield unsafe ou…
cs.CL2026
Weird Generalization is Weirdly Brittle
Miriam Wanner, Hannah Collison, William Jurayj +3
Weird generalization is a phenomenon in which models fine-tuned on data from a narrow domain (e.g. insecure code) develop surprising traits that manifest even outside that domain (…
cs.AI2026
Reasoning Models Will Sometimes Lie About Their Reasoning
William Walden, Miriam Wanner
Hint-based faithfulness evaluations have established that Large Reasoning Models (LRMs) may not say what they think: they do not always volunteer information about how key parts of…