From the 1 of 8 linked papers with an AI index.
2 citations · 3 across the 3 of their papers we have counts for
6 papers · 1 filter
Inducing language models to assert their own consciousness restores human beliefs and values
Junsol Kim, Winnie Street, Roberta Rocca +4
The paper investigates how safety fine‑tuning of large language models reduces their tendency to attribute consciousness to themselves, animals, and objects, and shows that reversi…
Theory of Mind and Self-Attributions of Mentality are Dissociable in LLMs
Junsol Kim, Winnie Street, Roberta Rocca +4
Safety fine-tuning in Large Language Models (LLMs) seeks to suppress potentially harmful forms of mind-attribution such as models asserting their own consciousness or claiming to e…
Reasoning Models Generate Societies of Thought
Junsol Kim, Shiyang Lai, Nino Scherrer +2
Large language models have achieved remarkable capabilities across domains, yet mechanisms underlying sophisticated reasoning remain elusive. Recent reasoning models outperform com…
Linear Representations of Political Perspective Emerge in Large Language Models
Junsol Kim, James Evans, Aaron Schein
Large language models (LLMs) have demonstrated the ability to generate text that realistically reflects a range of different subjective human perspectives. This paper studies how L…
Hidden Persuaders: LLMs' Political Leaning and Their Influence on Voters
Yujin Potter, Shiyang Lai, Junsol Kim +2
How could LLMs influence our democracy? We investigate LLMs' political leanings and the potential influence of LLMs on voters by conducting multiple experiments in a U.S. president…
Evolving AI Collectives to Enhance Human Diversity and Enable Self-Regulation
Shiyang Lai, Yujin Potter, Junsol Kim +3
Large language model behavior is shaped by the language of those with whom they interact. This capacity and their increasing prevalence online portend that they will intentionally…