3 papers
cs.CL2026
FindMyText: Robust, Scalable Detection of Text Containment in Large Web-Crawled Corpora
Lars Henry Berge Olsen, Pierre Lison, Martin Jullum +1
We present FindMyText, an open-source Python package designed to efficiently assess whether a given text appears, in part or in full, within a text corpus. The tool builds on prior…
cs.CL2025
Protecting De-identified Documents from Search-based Linkage Attacks
Pierre Lison, Mark Anderson
While de-identification models can help conceal the identity of the individuals mentioned in a document, they fail to address linkage risks, defined as the potential to map the de-…
cs.LG2025
Fairness-Aware Low-Rank Representation Fine-Tuning
Parameswaran Kamalaruban, Mark Anderson, Stuart Burrell +3
Pre-trained foundation models can be efficiently adapted for specific tasks using Low-Rank Adaptation (LoRA), but the fairness properties of these adapted classifiers remain undere…