29 citations · 44 across the 25 of their papers we have counts for
4 papers · 2 filters
Fixed Point Explainability
Emanuele La Malfa, Jon Vadillo, Marco Molinari +1
This paper introduces a formal notion of fixed point explanations, inspired by the "why regress" principle, to assess, through recursive applications, the stability of the interpla…
Out-of-Context Reasoning in Large Language Models
Jonathan Shaki, Emanuele La Malfa, Michael Wooldridge +1
We study how large language models (LLMs) reason about memorized knowledge through simple binary relations such as equality (), inequality (), and inclusion (). Unli…
Code Simulation as a Proxy for High-order Tasks in Large Language Models
Emanuele La Malfa, Christoph Weinhuber, Orazio Torre +6
Many reasoning, planning, and problem-solving tasks share an intrinsic algorithmic nature: correctly simulating each step is a sufficient condition to solve them correctly. We coll…
Jailbreaking Large Language Models in Infinitely Many Ways
Oliver Goldstein, Emanuele La Malfa, Felix Drinkall +2
We discuss the ``Infinitely Many Paraphrases'' attacks (IMP), a category of jailbreaks that leverages the increasing capabilities of a model to handle paraphrases and encoded commu…