4 citations · 4 across the 2 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It
Valen Tagliabue, Leonard Dung, Cameron Berg
Large language models sometimes behave in ways resembling human emotional responses, and recent work has identified internal representations that may explain this. We ask whether L…
cs.AI2025
Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of AI Welfare
Valen Tagliabue, Leonard Dung
We develop new experimental paradigms for measuring welfare in language models. We compare verbal reports of models about their preferences with preferences expressed through behav…