3 papers
cs.AI2026
Think Before You Lie: How Reasoning Leads to Honesty
Ann Yuan, Asma Ghandeharioun, Carter Blum +6
While existing evaluations of large language models (LLMs) measure deception rates, the underlying conditions that give rise to deceptive behavior are poorly understood. We investi…
cs.CL2025
Improving Neutral Point-of-View Generation with Data- and Parameter-Efficient RL
Jessica Hoffmann, Christiane Ahlheim, Zac Yu +8
The paper shows that parameter-efficient reinforcement learning (PE-RL) is a highly effective training regime to improve large language models' (LLMs) ability to answer queries on…
cs.CL2025
Beyond the Rosetta Stone: Unification Forces in Generalization Dynamics
Carter Blum, Katja Filippova, Ann Yuan +8
Large language models (LLMs) struggle with cross-lingual knowledge transfer: they hallucinate when asked in one language about facts expressed in a different language during traini…