4 papers
Data Provenance Auditing of Fine-Tuned Large Language Models with a Text-Preserving Technique
Yanming Li, Cédric Eichler, Nicolas Anciaux +3
We propose a system for marking sensitive or copyrighted texts to detect their use in fine-tuning large language models under black-box access with statistical guarantees. Our meth…
LPASS: Linear Probes as Stepping Stones for vulnerability detection using compressed LLMs
Luis Ibanez-Lissen, Lorena Gonzalez-Manzano, Jose Maria de Fuentes +1
Large Language Models (LLMs) are being extensively used for cybersecurity purposes. One of them is the detection of vulnerable codes. For the sake of efficiency and effectiveness,…
LUMIA: Linear probing for Unimodal and MultiModal Membership Inference Attacks leveraging internal LLM states
Luis Ibanez-Lissen, Lorena Gonzalez-Manzano, Jose Maria de Fuentes +2
Large Language Models (LLMs) are increasingly used in a variety of applications, but concerns around membership inference have grown in parallel. Previous efforts focus on black-to…
Nob-MIAs: Non-biased Membership Inference Attacks Assessment on Large Language Models with Ex-Post Dataset Construction
Cédric Eichler, Nathan Champeil, Nicolas Anciaux +3
The rise of Large Language Models (LLMs) has triggered legal and ethical concerns, especially regarding the unauthorized use of copyrighted materials in their training datasets. Th…