3 papers
cs.CR2026
Cordyceps: Covert Control Attacks on LLMs via Data Poisoning
Zedian Shao, Charles Fleming, Teodora Baluta
Large language models (LLMs) are often fine-tuned on uncurated text datasets that adversaries can poison. Existing poisoning attacks primarily rely on fixed trigger phrases that de…
cs.CR2025
Watermark Robustness and Radioactivity May Be at Odds in Federated Learning
Leixu Huang, Zedian Shao, Teodora Baluta
Federated learning (FL) enables fine-tuning large language models (LLMs) across distributed data sources. As these sources increasingly include LLM-generated text, provenance track…
cs.LG2024
Unlearning in- vs. out-of-distribution data in LLMs under gradient-based method
Teodora Baluta, Pascal Lamblin, Daniel Tarlow +2
Machine unlearning aims to solve the problem of removing the influence of selected training examples from a learned model. Despite the increasing attention to this problem, it rema…