Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Multilingual Unlearning in LLMs: Transfer, Dynamics, and Reversibility
Chaoyi Xiang, Olga Ohrimenko, Benjamin I. P. Rubinstein +1
Large language models (LLMs) can memorize sensitive facts, motivating unlearning methods that remove targeted knowledge without costly retraining. However, unlearning research rema…
cs.CL2025
TuBA: Cross-Lingual Transferability of Backdoor Attacks in LLMs with Instruction Tuning
Xuanli He, Jun Wang, Qiongkai Xu +4
The implications of backdoor attacks on English-centric large language models (LLMs) have been widely examined - such attacks can be achieved by embedding malicious behaviors durin…
cs.CL2024
SEEP: Training Dynamics Grounds Latent Representation Search for Mitigating Backdoor Poisoning Attacks
Xuanli He, Qiongkai Xu, Jun Wang +2
Modern NLP models are often trained on public datasets drawn from diverse sources, rendering them vulnerable to data poisoning attacks. These attacks can manipulate the model's beh…