activity
20242026
collaborators

7 papers

cs.LG2026

Training Large Language Models for Self-Explanation Faithfulness

Yeoktatt Cheah, María Pérez-Ortiz, Noah Y. Siegel +1

We propose a Reinforcement Learning (RL) method to directly optimize the faithfulness of self-explanations - the extent to which a model's generated reasoning accurately reflects i…

cs.CL2026

Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations

Noah Y. Siegel, Nicolas Heess, Maria Perez-Ortiz +1

When asked to explain their decisions, LLMs can often give explanations which sound plausible to humans. But are these explanations faithful, i.e. do they convey the factors actual…

cs.AI2026

A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior

Harry Mayne, Justin Singh Kang, Dewi Gould +3

LLM self-explanations are often presented as a promising tool for AI oversight, yet their faithfulness to the model's true reasoning process is poorly understood. Existing faithful…

cs.CR2025

LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring

Chloe Li, Mary Phuong, Noah Y. Siegel

Trustworthy evaluations of dangerous capabilities are increasingly crucial for determining whether an AI system is safe to deploy. One empirically demonstrated threat is sandbaggin…

cs.LG2025

Advancing Event Forecasting through Massive Training of Large Language Models: Challenges, Solutions, and Broader Impacts

Sang-Woo Lee, Sohee Yang, Donghyun Kwak +1

Many recent papers have studied the development of superforecaster-level event forecasting LLMs. While methodological problems with early studies cast doubt on the use of LLMs for…

cs.LG2024

On scalable oversight with weak LLMs judging strong LLMs

Zachary Kenton, Noah Y. Siegel, János Kramár +8

Scalable oversight protocols aim to enable humans to accurately supervise superhuman AI. In this paper we study debate, where two AI's compete to convince a judge; consultancy, whe…