Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
What Intermediate Layers Know: Detecting Jailbreaks from Entropy Dynamics
Sofiia Nikolenko, Michele Papucci, Mina Rezaei +1
Jailbreak attacks reveal a persistent weakness in aligned Large Language Models: carefully crafted prompts can elicit policy-violating responses despite safety training. While most…
cs.CL2025
Stress-testing Machine Generated Text Detection: Shifting Language Models Writing Style to Fool Detectors
Andrea Pedrotti, Michele Papucci, Cristiano Ciaccio +4
Recent advancements in Generative AI and Large Language Models (LLMs) have enabled the creation of highly realistic synthetic content, raising concerns about the potential for mali…