2 papers
cs.CL2026
Credal Large Language Models for Semantic Commitment under Uncertainty
Shireen Kudukkil Manchingal, Sofiia Nikolenko, Fabio Cuzzolin
Large language models (LLMs) often produce fluent but incorrect answers with unwarranted confidence. A central limitation is that standard LLMs represent uncertainty through a sing…
cs.CL2026
What Intermediate Layers Know: Detecting Jailbreaks from Entropy Dynamics
Sofiia Nikolenko, Michele Papucci, Mina Rezaei +1
Jailbreak attacks reveal a persistent weakness in aligned Large Language Models: carefully crafted prompts can elicit policy-violating responses despite safety training. While most…