works on

From the 1 of 6 linked papers with an AI index.

most citedInducing language models to assert their own consciousness restores human beliefs and values

2 citations · 2 across the 1 of their papers we have counts for

collaborators

6 papers

cs.CL20262 cited

Inducing language models to assert their own consciousness restores human beliefs and values

Junsol Kim, Winnie Street, Roberta Rocca +4

The paper investigates how safety fine‑tuning of large language models reduces their tendency to attribute consciousness to themselves, animals, and objects, and shows that reversi…

cs.CL2026

Theory of Mind and Self-Attributions of Mentality are Dissociable in LLMs

Junsol Kim, Winnie Street, Roberta Rocca +4

Safety fine-tuning in Large Language Models (LLMs) seeks to suppress potentially harmful forms of mind-attribution such as models asserting their own consciousness or claiming to e…

cs.CL2026

Reasoning Models Generate Societies of Thought

Junsol Kim, Shiyang Lai, Nino Scherrer +2

Large language models have achieved remarkable capabilities across domains, yet mechanisms underlying sophisticated reasoning remain elusive. Recent reasoning models outperform com…

cs.HC2025

Biased AI improves human decision-making but reduces trust

Shiyang Lai, Junsol Kim, Nadav Kunievsky +2

Current AI systems minimize risk by enforcing ideological neutrality, yet this may introduce automation bias by suppressing cognitive engagement in human decision-making. We conduc…

cs.CL2025

Linear Representations of Political Perspective Emerge in Large Language Models

Junsol Kim, James Evans, Aaron Schein

Large language models (LLMs) have demonstrated the ability to generate text that realistically reflects a range of different subjective human perspectives. This paper studies how L…

cs.CY2025

Differential impact from individual versus collective misinformation tagging on the diversity of Twitter (X) information engagement and mobility

Junsol Kim, Zhao Wang, Haohan Shi +2

Fears about the destabilizing impact of misinformation online have motivated individuals and platforms to respond. Individuals have increasingly challenged others' online claims wi…