Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Xeno-Interpretability: Investigating the Alien Minds of LLMs
F. Pierucci, M. Bracale Syrnikov, M. Prandi +3
Large language models are usually interpreted through concepts that humans already possess: truthfulness, refusal, deception, personality, harmfulness, and related categories. This…
cs.CL2026
Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety
Piercosma Bisconti, Matteo Prandi, Federico Pierucci +11
Background. Traditional safety benchmarks for language models evaluate generated text: whether a model outputs toxic language, reproduces bias, or follows harmful instructions. Whe…
cs.CL2026
From Adversarial Poetry to Adversarial Tales: An Interpretability Research Agenda
Piercosma Bisconti, Marcello Galisai, Matteo Prandi +6
Safety mechanisms in LLMs remain vulnerable to attacks that reframe harmful requests through culturally coded structures. We introduce Adversarial Tales, a jailbreak technique that…