2 papers
cs.CR2025
Improving Sustainability of Adversarial Examples in Class-Incremental Learning
Taifeng Liu, Xinjing Liu, Liangqiu Dong +3
Current adversarial examples (AEs) are typically designed for static models. However, with the wide application of Class-Incremental Learning (CIL), models are no longer static and…
cs.CL2025
Jinx: Unlimited LLMs for Probing Alignment Failures
Jiahao Zhao, Liwei Dong
Unlimited, or so-called helpful-only language models are trained without safety alignment constraints and never refuse user queries. They are widely used by leading AI companies as…