3 papers
cs.CR2026
Mechanistic Interpretability of LLM Jailbreaks via Internal Attribution Graphs
Anupam Wagle, Ifrat Ikhtear Uddin, Chaowei Zhang +1
Large language models (LLMs) exhibit remarkable capabilities but remain highly vulnerable to adversarial prompts and jailbreak attacks. Existing approaches primarily analyze these…
cs.CL2026
Acting Flatterers via LLMs Sycophancy: Combating Clickbait with LLMs Opposing-Stance Reasoning
Chaowei Zhang, Xiansheng Luo, Zewei Zhang +3
The widespread proliferation of online content has intensified concerns about clickbait, deceptive or exaggerated headlines designed to attract attention. While Large Language Mode…
cs.LG2026
Explainability-Guided Defense: Attribution-Aware Model Refinement Against Adversarial Data Attacks
Longwei Wang, Mohammad Navid Nayyem, Abdullah Al Rakin +3
The growing reliance on deep learning models in safety-critical domains such as healthcare and autonomous navigation underscores the need for defenses that are both robust to adver…