works on

From the 1 of 7 linked papers with an AI index.

collaborators

7 papers

cs.LG2026

ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions

Xinzhe Huang, Biwu Yao, Kedong Xiu +4

Recent research on Large Language Model (LLM) safety has widely adopted guardrails to identify unsafe LLM outputs. Existing guardrails typically formulate safety assessment as a de…

cs.LG2026

Confusion-Geometry Rebalancing for Long-Tailed Adversarial Training

Mengnan Zhao, Geyong Min, Lihe Zhang +2

Adversarial training under long tailed distributions suffers from a dual imbalance: the class imbalance skews the training objective toward head classes, and the adversarial inner…

cs.CR2026

AgentSnare: Learning to Delay, Divert, and Defuse Autonomous Penetration Agents

Ruoyu Wang, Heng Zhao, Renjie Wu +4

The paper presents AgentSnare, a system that dynamically creates deceptive decoy environments to mislead and delay autonomous penetration testing agents powered by large language m…

cs.CR2026

CoreUnlearn: Rethinking Concept Unlearning through Disentangled Component-Level Erasure in Text-guided Diffusion Models

Mengnan Zhao, Lihe Zhang, Baocai Yin

Text guided diffusion models have revolutionized image synthesis but also raise ethical concerns, such as privacy violation and harmful content generation. To mitigate these issues…

cs.LG2026

Unveiling the Backdoor Mechanism Hidden Behind Catastrophic Overfitting in Fast Adversarial Training

Mengnan Zhao, Lihe Zhang, Tianhang Zheng +2

Fast Adversarial Training (FAT) has attracted significant attention due to its efficiency in enhancing neural network robustness against adversarial attacks. However, FAT is prone…

cs.LG2026

Mitigating Error Amplification in Fast Adversarial Training

Mengnan Zhao, Lihe Zhang, Bo Wang +3

Fast Adversarial Training (FAT) has proven effective in enhancing model robustness by encouraging networks to learn perturbation-invariant representations. However, FAT often suffe…