3 papers
cs.CL2026
Does Out-of-Sight Equal Out-of-Mind in CoT Monitorability?
Pedro Ferreira, Wilker Aziz, Ivan Titov
Chain-of-thought (CoT) reasoning offers a window into the decision-making of large language models (LLMs), which can be monitored for target behaviors by reading the reasoning trac…
cs.CL2026
Truthful or Fabricated? Using Causal Attribution to Mitigate Reward Hacking in Explanations
Pedro Ferreira, Wilker Aziz, Ivan Titov
Chain-of-thought explanations are widely used to inspect the decision process of large language models (LLMs) and to evaluate the trustworthiness of model outputs, making them impo…
cs.CL2025
Explanation Regularisation through the Lens of Attributions
Pedro Ferreira, Ivan Titov, Wilker Aziz
Explanation regularisation (ER) has been introduced as a way to guide text classifiers to form their predictions relying on input tokens that humans consider plausible. This is ach…