From the 3 of 23 linked papers with an AI index.
4 papers · 1 filter
Metaphor Is Not All Attention Needs
Olga Sorokoletova, Francesco Giarrusso, Giacomo De Luca +6
Large language models are increasingly deployed in safety-critical applications, where their ability to resist harmful instructions is essential. Although post-training aims to mak…
Guarding the Guardrails: A Taxonomy-Driven Approach to Jailbreak Detection
Francesco Giarrusso, Olga E. Sorokoletova, Vincenzo Suriani +1
Jailbreaking techniques pose a significant threat to the safety of Large Language Models (LLMs). Existing defenses typically focus on single-turn attacks, lack coverage across lang…
From Adversarial Poetry to Adversarial Tales: An Interpretability Research Agenda
Piercosma Bisconti, Marcello Galisai, Matteo Prandi +6
Safety mechanisms in LLMs remain vulnerable to attacks that reframe harmful requests through culturally coded structures. We introduce Adversarial Tales, a jailbreak technique that…
Adversarial Poetry as a Universal Single-Turn Jailbreak Mechanism in Large Language Models
Piercosma Bisconti, Matteo Prandi, Federico Pierucci +7
We present evidence that adversarial poetry functions as a universal single-turn jailbreak technique for Large Language Models (LLMs). Across 25 frontier proprietary and open-weigh…