5 papers
Evaluating Concept Filtering Defenses against Child Sexual Abuse Material Generation by Text-to-Image Models
Ana-Maria Cretu, Klim Kireev, Amro Abdalla +5
We evaluate the effectiveness of filtering child images from training datasets of text-to-image models to prevent model misuse to create child sexual abuse material (CSAM). First,…
A Manually Annotated Image-Caption Dataset for Detecting Children in the Wild
Klim Kireev, Ana-Maria Creţu, Raphael Meier +3
Platforms and the law regulate digital content depicting minors (defined as individuals under 18 years of age) differently from other types of content. Given the sheer amount of co…
Characterizing and Detecting Propaganda-Spreading Accounts on Telegram
Klim Kireev, Yevhen Mykhno, Carmela Troncoso +1
Information-based attacks on social media, such as disinformation campaigns and propaganda, are emerging cybersecurity threats. The security community has focused on countering the…
SINBAD: Saliency-informed detection of breakage caused by ad blocking
Saiid El Hajj Chehade, Sandra Siby, Carmela Troncoso
Privacy-enhancing blocking tools based on filter-list rules tend to break legitimate functionality. Filter-list maintainers could benefit from automated breakage detection tools th…
Neural Exec: Learning (and Learning from) Execution Triggers for Prompt Injection Attacks
Dario Pasquini, Martin Strohmeier, Carmela Troncoso
We introduce a new family of prompt injection attacks, termed Neural Exec. Unlike known attacks that rely on handcrafted strings (e.g., "Ignore previous instructions and..."), we s…