From the 1 of 7 linked papers with an AI index.
7 papers
ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions
Xinzhe Huang, Biwu Yao, Kedong Xiu +4
Recent research on Large Language Model (LLM) safety has widely adopted guardrails to identify unsafe LLM outputs. Existing guardrails typically formulate safety assessment as a de…
Confusion-Geometry Rebalancing for Long-Tailed Adversarial Training
Mengnan Zhao, Geyong Min, Lihe Zhang +2
Adversarial training under long tailed distributions suffers from a dual imbalance: the class imbalance skews the training objective toward head classes, and the adversarial inner…
AgentSnare: Learning to Delay, Divert, and Defuse Autonomous Penetration Agents
Ruoyu Wang, Heng Zhao, Renjie Wu +4
The paper presents AgentSnare, a system that dynamically creates deceptive decoy environments to mislead and delay autonomous penetration testing agents powered by large language m…
CoreUnlearn: Rethinking Concept Unlearning through Disentangled Component-Level Erasure in Text-guided Diffusion Models
Mengnan Zhao, Lihe Zhang, Baocai Yin
Text guided diffusion models have revolutionized image synthesis but also raise ethical concerns, such as privacy violation and harmful content generation. To mitigate these issues…
Unveiling the Backdoor Mechanism Hidden Behind Catastrophic Overfitting in Fast Adversarial Training
Mengnan Zhao, Lihe Zhang, Tianhang Zheng +2
Fast Adversarial Training (FAT) has attracted significant attention due to its efficiency in enhancing neural network robustness against adversarial attacks. However, FAT is prone…
Mitigating Error Amplification in Fast Adversarial Training
Mengnan Zhao, Lihe Zhang, Bo Wang +3
Fast Adversarial Training (FAT) has proven effective in enhancing model robustness by encouraging networks to learn perturbation-invariant representations. However, FAT often suffe…