2 papers
cs.LG2026
ALERT: Zero-shot LLM Jailbreak Detection via Internal Discrepancy Amplification
Xiao Lin, Philip Li, Zhichen Zeng +6
Despite rich safety alignment strategies, large language models (LLMs) remain highly susceptible to jailbreak attacks, which compromise safety guardrails and pose serious security…
cs.CR2025
CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities
Yuxuan Zhu, Antony Kellermann, Dylan Bowman +13
Large language model (LLM) agents are increasingly capable of autonomously conducting cyberattacks, posing significant threats to existing applications. This growing risk highlight…