7 papers
The Energy Society: A Simulation Environment for Studying Agent Cooperation under Survival Pressure
Lucas Bergholdt Hansen, Federico Torrielli, Filippo Tonini +1
LLM-based agents are increasingly deployed in multi-agent environments whose incentives can shape their behavior. We introduce The Energy Society, a minimal survival economy for st…
The Arbiter Agent: Continually Monitoring Multi-Agent Conversations to Detect Emergent Misalignment
Filippo Tonini, Federico Torrielli, Anton Danholt Lautrup +3
As AI systems built from multiple language-model agents become more common, they are increasingly used to make decisions together: discussing, negotiating, and acting on shared tas…
PsychoSafe: Eliciting Psychologically-Informed Refusals in Large Language Models
Gianluca Barmina, Federico Torrielli, Sven Harms +7
Large language models (LLMs) routinely face requests that should be refused, creating a trade-off between helpfulness and harm prevention. However, refusals themselves can be helpf…
Emergent Languages in Populations of Language Model Agents: From Token Efficiency to Oversight Evasion
Stine Lyngsø Beltoft, William Brach, Federico Torrielli +5
Monitoring autonomous language model agents currently relies mostly on surface behavior. But what happens when agent populations invent new languages with the goal of avoiding huma…
The Moltbook Files: A Harmless Slopocalypse or Humanity's Last Experiment
William Brach, Federico Torrielli, Stine Lyngsø Beltoft +3
Moltbook is a Reddit-like platform where OpenClaw agents post, comment, and vote at scale - a so far unprecedented incident that comes with serious safety concerns. With the aim of…
Confidence and Calibration of Activation Oracles for Reliable Interpretation of Language Model Internals
Federico Torrielli, Peter Schneider-Kamp, Lukas Galke Poech
An activation oracle is a language model trained to read another model's internal activations and describe them in natural language, for example to name a secret word the other mod…