most citedInstitutional AI: A Governance Framework for Distributional AGI Safety

2 citations · 2 across the 6 of their papers we have counts for

collaborators

7 papers

cs.CY2026

Standards for trustworthy AI in the European Union: technical rationale, structural challenges, and an implementation path

Piercosma Bisconti, Marcello Galisai

This white paper examines the technical foundations of European AI standardization under the AI Act. It explains how harmonized standards enable the presumption of conformity mecha…

cs.GT2026

Institutional AI: Governing LLM Collusion in Multi-Agent Cournot Markets via Public Governance Graphs

Marcantonio Bracale Syrnikov, Federico Pierucci, Marcello Galisai +6

Multi-agent LLM ensembles can converge on coordinated, socially harmful equilibria. This paper advances an experimental framework for evaluating Institutional AI, our system-level…

cs.CY20262 cited

Institutional AI: A Governance Framework for Distributional AGI Safety

Federico Pierucci, Marcello Galisai, Marcantonio Syrnikov Bracale +6

As LLM-based systems increasingly operate as agents embedded within human social and technical systems, alignment can no longer be treated as a property of an isolated model, but m…

cs.CL2026

From Adversarial Poetry to Adversarial Tales: An Interpretability Research Agenda

Piercosma Bisconti, Marcello Galisai, Matteo Prandi +6

Safety mechanisms in LLMs remain vulnerable to attacks that reframe harmful requests through culturally coded structures. We introduce Adversarial Tales, a jailbreak technique that…

cs.MA2025

Beyond Single-Agent Safety: A Taxonomy of Risks in LLM-to-LLM Interactions

Piercosma Bisconti, Marcello Galisai, Federico Pierucci +2

This paper examines why safety mechanisms designed for human-model interaction do not scale to environments where large language models (LLMs) interact with each other. Most curren…

cs.CL2025

Adversarial Poetry as a Universal Single-Turn Jailbreak Mechanism in Large Language Models

Piercosma Bisconti, Matteo Prandi, Federico Pierucci +7

We present evidence that adversarial poetry functions as a universal single-turn jailbreak technique for Large Language Models (LLMs). Across 25 frontier proprietary and open-weigh…