13 papers
Draining the Energy Commons: Self-Defeating Over-Appropriation as a Coordination Failure in Agentic LLM Collectives
Marcantonio Bracale Syrnicov, Federico Pierucci, Matteo Prandi +4
LLMs are increasingly deployed as agents that plan, use tools, and act over time. When they share persistent resources, such as compute pools or energy reserves, decisions by one a…
Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety
Piercosma Bisconti, Matteo Prandi, Federico Pierucci +11
Background. Traditional safety benchmarks for language models evaluate generated text: whether a model outputs toxic language, reproduces bias, or follows harmful instructions. Whe…
Metaphor Is Not All Attention Needs
Olga Sorokoletova, Francesco Giarrusso, Giacomo De Luca +6
Large language models are increasingly deployed in safety-critical applications, where their ability to resist harmful instructions is essential. Although post-training aims to mak…
Adversarial Humanities Benchmark: Results on Stylistic Robustness in Frontier Model Safety
Marcello Galisai, Susanna Cifani, Francesco Giarrusso +5
The Adversarial Humanities Benchmark (AHB) evaluates whether model safety refusals survive a shift away from familiar harmful prompt forms. Starting from harmful tasks drawn from M…
Agentic Microphysics: A Manifesto for Generative AI Safety
Federico Pierucci, Matteo Prandi, Marcantonio Bracale Syrnikov +2
This paper advances a methodological proposal for safety research in agentic AI. As systems acquire planning, memory, tool use, persistent identity, and sustained interaction, safe…
AI Agents Under EU Law
Luca Nannini, Adam Leon Smith, Michele Joshua Maggini +6
AI agents - i.e. AI systems that autonomously plan, invoke external tools, and execute multi-step action chains with reduced human involvement - are being deployed at scale across…