natural language processing

Pretraining Data Can Be Poisoned through Computational Propaganda

arXiv:2607.15267

summary

The paper shows that language model pretraining data can be poisoned through publicly editable web discussion pages, and introduces a method called HalfLife to estimate how much malicious content survives web crawling and curation.

Abstract

Poisoning pretraining data can introduce harmful behaviors to LMs that are difficult to detect and mitigate. Prior work on poisoning pretraining data has largely exploited established data sources such as Wikipedia, which do not represent the large scale and heterogeneity typical of pretraining corpora, and has ignored the interaction between poisoned data and data curation pipelines. We demonstrate that poisoning attacks on pretraining data are feasible beyond this limited setting through an existing web-scale content injection mechanism: public discussion interfaces. Additionally, to measure whether malicious content is included after web crawling and data curation, we introduce HalfLife, a novel analysis for estimating adversarial content inclusion in web-crawl based LM training data. We use HalfLife to explore the feasibility of poisoning pretraining corpora at web scale through open discussion interfaces. Our analysis demonstrates the importance of estimating whether poison injections are included in pretraining data, and establishes third-party webpage content as a possible vector for attacking language model pretraining.

Topics & keywords

#data poisoning#language model pretraining#web crawling#adversarial attacks#content injectionpretraining data poisoningHalfLife analysispublic discussion interfacesdata curation pipelinesweb-scale corpus
Pretraining Data Can Be Poisoned through Computational Propaganda · wovepaper