Showing cs.CRShow all
2 papers · 1 filter
cs.CR2026
SafeRedirect: Defeating Internal Safety Collapse via Task-Completion Redirection in Frontier LLMs
Chao Pan, Yu Wu, Xin Yao
Internal Safety Collapse (ISC) is a failure mode in which frontier LLMs, when executing legitimate professional tasks whose correct completion structurally requires harmful content…
cs.CR2024
Joint Universal Adversarial Perturbations with Interpretations
Liang-bo Ning, Zeyu Dai, Wenqi Fan +4
Deep neural networks (DNNs) have significantly boosted the performance of many challenging tasks. Despite the great development, DNNs have also exposed their vulnerability. Recent…