Showing cs.CRShow all
3 papers · 1 filter
cs.CR2026
Stop Tracking Me! Proactive Defense Against Attribute Inference Attack in LLMs
Dong Yan, Jian Liang, Ran He +1
Recent studies have shown that large language models (LLMs) can infer private user attributes (e.g., age, location, gender) from user-generated text shared online, enabling rapid a…
cs.CR2025
Do We Really Need Curated Malicious Data for Safety Alignment in Multi-modal Large Language Models?
Yanbo Wang, Jiyang Guan, Jian Liang +1
Multi-modal large language models (MLLMs) have made significant progress, yet their safety alignment remains limited. Typically, current open-source MLLMs rely on the alignment inh…
cs.CR2025
Protecting Model Adaptation from Trojans in the Unlabeled Data
Lijun Sheng, Jian Liang, Ran He +2
Model adaptation tackles the distribution shift problem with a pre-trained model instead of raw data, which has become a popular paradigm due to its great privacy protection. Exist…