3 papers
cs.CL2025
Hallucination Detection via Internal States and Structured Reasoning Consistency in Large Language Models
Yusheng Song, Lirong Qiu, Xi Zhang +1
The detection of sophisticated hallucinations in Large Language Models (LLMs) is hampered by a ``Detection Dilemma'': methods probing internal states (Internal State Probing) excel…
cs.AI2025
Beyond Surface-Level Detection: Towards Cognitive-Driven Defense Against Jailbreak Attacks via Meta-Operations Reasoning
Rui Pu, Chaozhuo Li, Rui Ha +3
Defending large language models (LLMs) against jailbreak attacks is essential for their safe and reliable deployment. Existing defenses often rely on shallow pattern matching, whic…
cs.CR2025
MirrorShield: Towards Universal Defense Against Jailbreaks via Entropy-Guided Mirror Crafting
Rui Pu, Chaozhuo Li, Rui Ha +3
Defending large language models (LLMs) against jailbreak attacks is crucial for ensuring their safe deployment. Existing defense strategies typically rely on predefined static crit…