2 papers
cs.LG2026
Why Do Language Model Agents Whistleblow?
Kushal Agrawal, Frank Xiao, Guido Bergman +1
The deployment of Large Language Models (LLMs) as tool-using agents causes their alignment training to manifest in new ways. Recent work finds that language models can use tools in…
cs.CL2025
AI Debaters are More Persuasive when Arguing in Alignment with Their Own Beliefs
MarÃa Victoria Carro, Denise Alejandra Mester, Facundo Nieto +9
The core premise of AI debate as a scalable oversight technique is that it is harder to lie convincingly than to refute a lie, enabling the judge to identify the correct position.…