2 papers
cs.LG2025
Tokens for Learning, Tokens for Unlearning: Mitigating Membership Inference Attacks in Large Language Models via Dual-Purpose Training
Toan Tran, Ruixuan Liu, Li Xiong
Large language models (LLMs) have become the backbone of modern natural language processing but pose privacy concerns about leaking sensitive training data. Membership inference at…
cs.CR2024
ExpShield: Safeguarding Web Text from Unauthorized Crawling and LLM Exploitation
Ruixuan Liu, Toan Tran, Tianhao Wang +3
As large language models increasingly memorize web-scraped training content, they risk exposing copyrighted or private information. Existing protections require compliance from cra…