5 papers
PsychoSafe: Eliciting Psychologically-Informed Refusals in Large Language Models
Gianluca Barmina, Federico Torrielli, Sven Harms +7
Large language models (LLMs) routinely face requests that should be refused, creating a trade-off between helpfulness and harm prevention. However, refusals themselves can be helpf…
Emergent Languages in Populations of Language Model Agents: From Token Efficiency to Oversight Evasion
Stine Lyngsø Beltoft, William Brach, Federico Torrielli +5
Monitoring autonomous language model agents currently relies mostly on surface behavior. But what happens when agent populations invent new languages with the goal of avoiding huma…
The Moltbook Files: A Harmless Slopocalypse or Humanity's Last Experiment
William Brach, Federico Torrielli, Stine Lyngsø Beltoft +3
Moltbook is a Reddit-like platform where OpenClaw agents post, comment, and vote at scale - a so far unprecedented incident that comes with serious safety concerns. With the aim of…
SDUs DAISY: A Benchmark for Danish Culture
Jacob Nielsen, Stine L. Beltoft, Peter Schneider-Kamp +1
We introduce Daisy, a factual knowledge benchmark for Danish cultural heritage, based on curated topics from the Danish Culture Canon 2006. For each artifact in the culture canon,…
Not Everything That Counts Can Be Counted: A Case for Safe Qualitative AI
Stine Beltoft, Lukas Galke
Artificial intelligence (AI) and large language models (LLM) are reshaping science, with most recent advances culminating in fully-automated scientific discovery pipelines. But qua…