3 papers
cs.CL2025
LexiMark: Robust Watermarking via Lexical Substitutions to Enhance Membership Verification of an LLM's Textual Training Data
Eyal German, Sagiv Antebi, Edan Habler +2
Large language models (LLMs) can be trained or fine-tuned on data obtained without the owner's consent. Verifying whether a specific LLM was trained on particular data instances or…
cs.CR2025
Tag&Tab: Pretraining Data Detection in Large Language Models Using Keyword-Based Membership Inference Attack
Sagiv Antebi, Edan Habler, Asaf Shabtai +1
Large language models (LLMs) have become essential tools for digital task assistance. Their training relies heavily on the collection of vast amounts of data, which may include cop…
cs.CR2025
Tab-MIA: A Benchmark Dataset for Membership Inference Attacks on Tabular Data in LLMs
Eyal German, Sagiv Antebi, Daniel Samira +2
Large language models (LLMs) are increasingly trained on tabular data, which, unlike unstructured text, often contains personally identifiable information (PII) in a highly structu…