2 papers
cs.CL2025
AI Debaters are More Persuasive when Arguing in Alignment with Their Own Beliefs
María Victoria Carro, Denise Alejandra Mester, Facundo Nieto +9
The core premise of AI debate as a scalable oversight technique is that it is harder to lie convincingly than to refute a lie, enabling the judge to identify the correct position.…
cs.LG2025
Why Do Language Model Agents Whistleblow?
Kushal Agrawal, Frank Xiao, Guido Bergman +1
The deployment of Large Language Models (LLMs) as tool-using agents causes their alignment training to manifest in new ways. Recent work finds that language models can use tools in…