From the 1 of 3 linked papers with an AI index.
4 papers · 1 filter
FindMyText: Robust, Scalable Detection of Text Containment in Large Web-Crawled Corpora
Lars Henry Berge Olsen, Pierre Lison, Martin Jullum +1
FindMyText is an open‑source Python tool that efficiently checks whether a given text appears, fully or partially, in large web‑crawled corpora by using fingerprint chains to detec…
Protecting De-identified Documents from Search-based Linkage Attacks
Pierre Lison, Mark Anderson
While de-identification models can help conceal the identity of the individuals mentioned in a document, they fail to address linkage risks, defined as the potential to map the de-…
Truthful Text Sanitization Guided by Inference Attacks
Ildikó Pilán, Benet Manzanares-Salor, David Sánchez +1
Text sanitization aims to rewrite parts of a document to prevent disclosure of personal information. The central challenge of text sanitization is to strike a balance between priva…
Conversational Feedback in Scripted versus Spontaneous Dialogues: A Comparative Analysis
Ildikó Pilán, Laurent Prévot, Hendrik Buschmeier +1
Scripted dialogues such as movie and TV subtitles constitute a widespread source of training data for conversational NLP models. However, there are notable linguistic differences b…