2 citations · 2 across the 5 of their papers we have counts for
7 papers
Institutional AI: Governing LLM Collusion in Multi-Agent Cournot Markets via Public Governance Graphs
Marcantonio Bracale Syrnikov, Federico Pierucci, Marcello Galisai +6
Multi-agent LLM ensembles can converge on coordinated, socially harmful equilibria. This paper advances an experimental framework for evaluating Institutional AI, our system-level…
Institutional AI: A Governance Framework for Distributional AGI Safety
Federico Pierucci, Marcello Galisai, Marcantonio Syrnikov Bracale +6
As LLM-based systems increasingly operate as agents embedded within human social and technical systems, alignment can no longer be treated as a property of an isolated model, but m…
From Adversarial Poetry to Adversarial Tales: An Interpretability Research Agenda
Piercosma Bisconti, Marcello Galisai, Matteo Prandi +6
Safety mechanisms in LLMs remain vulnerable to attacks that reframe harmful requests through culturally coded structures. We introduce Adversarial Tales, a jailbreak technique that…
Adversarial Poetry as a Universal Single-Turn Jailbreak Mechanism in Large Language Models
Piercosma Bisconti, Matteo Prandi, Federico Pierucci +7
We present evidence that adversarial poetry functions as a universal single-turn jailbreak technique for Large Language Models (LLMs). Across 25 frontier proprietary and open-weigh…
Guarding the Guardrails: A Taxonomy-Driven Approach to Jailbreak Detection
Francesco Giarrusso, Olga E. Sorokoletova, Vincenzo Suriani +1
Jailbreaking techniques pose a significant threat to the safety of Large Language Models (LLMs). Existing defenses typically focus on single-turn attacks, lack coverage across lang…
Towards a scalable AI-driven framework for data-independent Cyber Threat Intelligence Information Extraction
Olga Sorokoletova, Emanuele Antonioni, Giordano Colò
Cyber Threat Intelligence (CTI) is critical for mitigating threats to organizations, governments, and institutions, yet the necessary data are often dispersed across diverse format…