3 papers
cs.CL2025
Sanitize Your Responses: Mitigating Privacy Leakage in Large Language Models
Wenjie Fu, Huandong Wang, Junyao Gao +2
As Large Language Models (LLMs) achieve remarkable success across a wide range of applications, such as chatbots and code copilots, concerns surrounding the generation of harmful c…
cs.AI2025
A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy
Huandong Wang, Wenjie Fu, Yingzhou Tang +7
While large language models (LLMs) present significant potential for supporting numerous real-world applications and delivering positive social impacts, they still face significant…
cs.CL2024
Practical Membership Inference Attacks against Fine-tuned Large Language Models via Self-prompt Calibration
Wenjie Fu, Huandong Wang, Chen Gao +3
Membership Inference Attacks (MIA) aim to infer whether a target data record has been utilized for model training or not. Existing MIAs designed for large language models (LLMs) ca…