1 citations · 1 across the 2 of their papers we have counts for
7 papers
Conversational AI increases political knowledge as effectively as self-directed internet search
Lennart Luettgau, Hannah Rose Kirk, Kobi Hackenburg +6
Conversational AI systems are increasingly being used in place of traditional search engines to help users complete information-seeking tasks. This has raised concerns in the polit…
People readily follow personal advice from AI but it does not improve their well-being
Lennart Luettgau, Vanessa Cheung, Magda Dubois +8
People increasingly seek personal advice from large language models (LLMs), yet whether humans follow their advice, and its consequences for their well-being, remains unknown. In a…
Reward Models Inherit Value Biases from Pretraining
Brian Christian, Jessica A. F. Thompson, Elle Michelle Yang +4
Reward models (RMs) are central to aligning large language models (LLMs) with human values but have received less attention than pretrained and post-trained LLMs themselves. Becaus…
Neural steering vectors reveal dose and exposure-dependent impacts of human-AI relationships
Hannah Rose Kirk, Henry Davidson, Ed Saunders +4
Humans are increasingly forming parasocial relationships with AI systems, and modern AI shows an increasing tendency to display social and relationship-seeking behaviour. However,…
Reward Model Interpretability via Optimal and Pessimal Tokens
Brian Christian, Hannah Rose Kirk, Jessica A. F. Thompson +2
Reward modeling has emerged as a crucial component in aligning large language models with human values. Significant attention has focused on using reward models as a means for fine…
Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language
Christopher Summerfield, Lennart Luettgau, Magda Dubois +9
We examine recent research that asks whether current AI systems may be developing a capacity for "scheming" (covertly and strategically pursuing misaligned goals). We compare curre…