4 papers
Secure Forgetting: A Framework for Privacy-Driven Unlearning in Large Language Model (LLM)-Based Agents
Dayong Ye, Tainqing Zhu, Congcong Zhu +5
Large language model (LLM)-based agents have recently gained considerable attention due to the powerful reasoning capabilities of LLMs. Existing research predominantly focuses on e…
Shapley-Guided Neural Repair Approach via Derivative-Free Optimization
Xinyu Sun, Wanwei Liu, Haoang Chi +7
DNNs are susceptible to defects like backdoors, adversarial attacks, and unfairness, undermining their reliability. Existing approaches mainly involve retraining, optimization, con…
Turning Black Box into White Box: Dataset Distillation Leaks
Huajie Chen, Tianqing Zhu, Yuchen Zhong +7
Dataset distillation compresses a large real dataset into a small synthetic one, enabling models trained on the synthetic data to achieve performance comparable to those trained on…
Isolate Trigger: Detecting and Eliminating Adaptive Backdoor Attacks
Chengrui Sun, Hua Zhang, Haoran Gao +7
Deep learning models are widely deployed in various applications but remain vulnerable to stealthy adversarial threats, particularly backdoor attacks. Backdoor models trained on po…