1 citations · 1 across the 8 of their papers we have counts for
3 papers · 1 filter
Credal Large Language Models for Semantic Commitment under Uncertainty
Shireen Kudukkil Manchingal, Sofiia Nikolenko, Fabio Cuzzolin
Large language models (LLMs) often produce fluent but incorrect answers with unwarranted confidence. A central limitation is that standard LLMs represent uncertainty through a sing…
What Intermediate Layers Know: Detecting Jailbreaks from Entropy Dynamics
Sofiia Nikolenko, Michele Papucci, Mina Rezaei +1
Jailbreak attacks reveal a persistent weakness in aligned Large Language Models: carefully crafted prompts can elicit policy-violating responses despite safety training. While most…
Random-Set Large Language Models
Muhammad Mubashar, Shireen Kudukkil Manchingal, Fabio Cuzzolin
Large Language Models (LLMs) are known to produce very high-quality tests and responses to our queries. But how much can we trust this generated text? In this paper, we study the p…