2 citations · 2 across the 6 of their papers we have counts for
7 papers
Standards for trustworthy AI in the European Union: technical rationale, structural challenges, and an implementation path
Piercosma Bisconti, Marcello Galisai
This white paper examines the technical foundations of European AI standardization under the AI Act. It explains how harmonized standards enable the presumption of conformity mecha…
Institutional AI: Governing LLM Collusion in Multi-Agent Cournot Markets via Public Governance Graphs
Marcantonio Bracale Syrnikov, Federico Pierucci, Marcello Galisai +6
Multi-agent LLM ensembles can converge on coordinated, socially harmful equilibria. This paper advances an experimental framework for evaluating Institutional AI, our system-level…
Institutional AI: A Governance Framework for Distributional AGI Safety
Federico Pierucci, Marcello Galisai, Marcantonio Syrnikov Bracale +6
As LLM-based systems increasingly operate as agents embedded within human social and technical systems, alignment can no longer be treated as a property of an isolated model, but m…
From Adversarial Poetry to Adversarial Tales: An Interpretability Research Agenda
Piercosma Bisconti, Marcello Galisai, Matteo Prandi +6
Safety mechanisms in LLMs remain vulnerable to attacks that reframe harmful requests through culturally coded structures. We introduce Adversarial Tales, a jailbreak technique that…
Beyond Single-Agent Safety: A Taxonomy of Risks in LLM-to-LLM Interactions
Piercosma Bisconti, Marcello Galisai, Federico Pierucci +2
This paper examines why safety mechanisms designed for human-model interaction do not scale to environments where large language models (LLMs) interact with each other. Most curren…
Adversarial Poetry as a Universal Single-Turn Jailbreak Mechanism in Large Language Models
Piercosma Bisconti, Matteo Prandi, Federico Pierucci +7
We present evidence that adversarial poetry functions as a universal single-turn jailbreak technique for Large Language Models (LLMs). Across 25 frontier proprietary and open-weigh…