papers

Publications (187)

cs.LG2026

Pruning Unsafe Tickets: A Resource-Efficient Framework for Safer and More Robust LLMs

Wai Man Si, Mingjie Li, Michael Backes +1

Machine learning models are increasingly deployed in real-world applications, but even aligned models such as Mistral and LLaVA still exhibit unsafe behaviors inherited from pre-tr…

cs.CR2024

Image-Perfect Imperfections: Safety, Bias, and Authenticity in the Shadow of Text-To-Image Model Evolution

Yixin Wu, Yun Shen, Michael Backes +1

Text-to-image models, such as Stable Diffusion (SD), undergo iterative updates to improve image quality and address concerns such as safety. Improvements in image quality are strai…

cs.CL2026

SafeReview: Defending LLM-based Review Systems Against Adversarial Hidden Prompts

Yuan Xin, Yixuan Weng, Minjun Zhu +5

As Large Language Models (LLMs) are increasingly integrated into academic peer review, their vulnerability to adversarial hidden prompts, i.e., adversarial instructions embedded in…

astro-ph.HE2023

Redshift determination of blazars for the Cherenkov Telescope Array

Eli Kasai, Paolo Goldoni, Santiago Pita +5

Blazars are the most numerous type of observed high-energy gamma-ray emitters. However, their emission mechanisms and population properties are still not well-understood. Crucial t…

cs.CR2020

Decentralized Privacy-Preserving Proximity Tracing

Carmela Troncoso, Mathias Payer, Jean-Pierre Hubaux +31

This document describes and analyzes a system for secure and privacy-preserving proximity tracing at large scale. This system, referred to as DP3T, provides a technological foundat…

astro-ph.HE2024

Chasing Gravitational Waves with the Cherenkov Telescope Array

Jarred Gershon Green, Alessandro Carosi, Lara Nava +563

The detection of gravitational waves from a binary neutron star merger by Advanced LIGO and Advanced Virgo (GW170817), along with the discovery of the electromagnetic counterparts…

cs.LG2025

Synthetic Artifact Auditing: Tracing LLM-Generated Synthetic Data Usage in Downstream Applications

Yixin Wu, Ziqing Yang, Yun Shen +2

Large language models (LLMs) have facilitated the generation of high-quality, cost-effective synthetic data for developing downstream models and conducting statistical analyses in…

cs.CR2019

Updates-Leak: Data Set Inference and Reconstruction Attacks in Online Learning

Ahmed Salem, Apratim Bhattacharya, Michael Backes +2

Machine learning (ML) has progressed rapidly during the past decade and the major factor that drives such development is the unprecedented large-scale data. As data generation is a…

cs.CR2017

walk2friends: Inferring Social Links from Mobility Profiles

Michael Backes, Mathias Humbert, Jun Pang +1

The development of positioning technologies has resulted in an increasing amount of mobility data being available. While bringing a lot of convenience to people's life, such availa…

cs.AI2025

Are We in the AI-Generated Text World Already? Quantifying and Monitoring AIGT on Social Media

Zhen Sun, Zongmin Zhang, Xinyue Shen +5

Social media platforms are experiencing a growing presence of AI-Generated Texts (AIGTs). However, the misuse of AIGTs could have profound implications for public opinion, such as…

cs.CR2017

Stack Overflow Considered Harmful? The Impact of Copy&Paste on Android Application Security

Felix Fischer, Konstantin Böttinger, Huang Xiao +4

Online programming discussion platforms such as Stack Overflow serve as a rich source of information for software developers. Available information include vibrant discussions and…

cs.CR2019

Automated Verification of Accountability in Security Protocols

Robert Künnemann, Ilkan Esiyok, Michael Backes

Accountability is a recent paradigm in security protocol design which aims to eliminate traditional trust assumptions on parties and hold them accountable for their misbehavior. It…

cs.CR2026

Benchmark of Benchmarks: Unpacking Influence and Code Repository Quality in LLM Safety Benchmarks

Junjie Chu, Xinyue Shen, Ye Leng +3

The rapid expansion of research in LLM safety presents challenges in tracking advancements, making benchmarks important evaluation infrastructures for identifying key trends and fa…

cs.SI2022

On Xing Tian and the Perseverance of Anti-China Sentiment Online

Xinyue Shen, Xinlei He, Michael Backes +3

Sinophobia, anti-Chinese sentiment, has existed on the Web for a long time. The outbreak of COVID-19 and the extended quarantine has further amplified it. However, we lack a quanti…

cs.CR2022

Backdoor Attacks in the Supply Chain of Masked Image Modeling

Xinyue Shen, Xinlei He, Zheng Li +3

Masked image modeling (MIM) revolutionizes self-supervised learning (SSL) for image pre-training. In contrast to previous dominating self-supervised methods, i.e., contrastive lear…

cs.CR2020

Adversarial Attacks on Classifiers for Eye-based User Modelling

Inken Hagestedt, Michael Backes, Andreas Bulling

An ever-growing body of work has demonstrated the rich information content available in eye movements for user modelling, e.g. for predicting users' activities, cognitive processes…

cs.CR2017

Deemon: Detecting CSRF with Dynamic Analysis and Property Graphs

Giancarlo Pellegrino, Martin Johns, Simon Koch +2

Cross-Site Request Forgery (CSRF) vulnerabilities are a severe class of web vulnerabilities that have received only marginal attention from the research and security testing commun…

cs.CY2026

On the Trustworthiness of Generative Foundation Models: Guideline, Assessment, and Perspective

Yue Huang, Chujie Gao, Siyuan Wu +63

Generative Foundation Models (GenFMs) have emerged as transformative tools. However, their widespread adoption raises critical concerns regarding trustworthiness across dimensions.…

cs.LG2025

Fairness and/or Privacy on Social Graphs

Bartlomiej Surma, Michael Backes, Yang Zhang

Graph Neural Networks (GNNs) have shown remarkable success in various graph-based learning tasks. However, recent studies have raised concerns about fairness and privacy issues in…

cs.CR2020

Stealing Links from Graph Neural Networks

Xinlei He, Jinyuan Jia, Michael Backes +2

Graph data, such as chemical networks and social networks, may be deemed confidential/private because the data owner often spends lots of resources collecting the data or the data…

cs.CR2019

Adversarial Vulnerability Bounds for Gaussian Process Classification

Michael Thomas Smith, Kathrin Grosse, Michael Backes +1

Machine learning (ML) classification is increasingly used in safety-critical systems. Protecting ML classifiers from adversarial examples is crucial. We propose that the main threa…

cs.CR2025

Adjacent Words, Divergent Intents: Jailbreaking Large Language Models via Task Concurrency

Yukun Jiang, Mingjie Li, Michael Backes +1

Despite their superior performance on a wide range of domains, large language models (LLMs) remain vulnerable to misuse for generating harmful content, a risk that has been further…

cs.CR2025

Excessive Reasoning Attack on Reasoning LLMs

Wai Man Si, Mingjie Li, Michael Backes +1

Recent reasoning large language models (LLMs), such as OpenAI o1 and DeepSeek-R1, exhibit strong performance on complex tasks through test-time inference scaling. However, prior st…

cs.CR2026

Real Money, Fake Models: Deceptive Model Claims in Shadow APIs

Yage Zhang, Yukun Jiang, Zeyuan Chen +3

Access to frontier large language models (LLMs), such as GPT-5 and Gemini-2.5, is often hindered by high pricing, payment barriers, and regional restrictions. These limitations dri…

cs.CR2016

ARTist: The Android Runtime Instrumentation and Security Toolkit

Michael Backes, Sven Bugiel, Oliver Schranz +2

We present ARTist, a compiler-based application instrumentation solution for Android. ARTist is based on the new ART runtime and the on-device dex2oat compiler of Android, which re…

cs.CR2023

Prompt Backdoors in Visual Prompt Learning

Hai Huang, Zhengyu Zhao, Michael Backes +2

Fine-tuning large pre-trained computer vision models is infeasible for resource-limited users. Visual prompt learning (VPL) has thus emerged to provide an efficient and flexible al…

cs.CV2026

MultiMem: Measuring and Mitigating Memorization in Multi-Modal Contrastive Learning

Wenhao Wang, Franziska Boenisch, Michael Backes +1

Memorization in machine learning models enables high performance on rare in-distribution samples by capturing their atypical patterns. However, it also causes harmful retention of…

cs.CR2018

ML-Leaks: Model and Data Independent Membership Inference Attacks and Defenses on Machine Learning Models

Ahmed Salem, Yang Zhang, Mathias Humbert +3

Machine learning (ML) has become a core component of many real-world applications and training data is a key factor that drives current progress. This huge success has led Internet…

cs.CR2021

Inference Attacks Against Graph Neural Networks

Zhikun Zhang, Min Chen, Michael Backes +2

Graph is an important data representation ubiquitously existing in the real world. However, analyzing the graph data is computationally difficult due to its non-Euclidean nature. G…

cs.CR2019

The Limitations of Model Uncertainty in Adversarial Settings

Kathrin Grosse, David Pfaff, Michael Thomas Smith +1

Machine learning models are vulnerable to adversarial examples: minor perturbations to input samples intended to deliberately cause misclassification. While an obvious security thr…

cs.CR2023

FAKEPCD: Fake Point Cloud Detection via Source Attribution

Yiting Qu, Zhikun Zhang, Yun Shen +2

To prevent the mischievous use of synthetic (fake) point clouds produced by generative models, we pioneer the study of detecting point cloud authenticity and attributing them to th…

cs.CR2016

From Closed-world Enforcement to Open-world Assessment of Privacy

Michael Backes, Pascal Berrang, Praveen Manoharan

In this paper, we develop a user-centric privacy framework for quantitatively assessing the exposure of personal information in open settings. Our formalization addresses key-chall…

cs.CL2024

TrustLLM: Trustworthiness in Large Language Models

Yue Huang, Lichao Sun, Haoran Wang +67

Large language models (LLMs), exemplified by ChatGPT, have gained considerable attention for their excellent natural language processing capabilities. Nonetheless, these LLMs prese…

cs.CR2022

Auditing Membership Leakages of Multi-Exit Networks

Zheng Li, Yiyong Liu, Xinlei He +3

Relying on the fact that not all inputs require the same amount of computation to yield a confident prediction, multi-exit networks are gaining attention as a prominent approach fo…

cs.CR2023

Backdoor Attacks Against Dataset Distillation

Yugeng Liu, Zheng Li, Michael Backes +2

Dataset distillation has emerged as a prominent technique to improve data efficiency when training machine learning models. It encapsulates the knowledge from a large dataset into…

cs.CR2024

Link Stealing Attacks Against Inductive Graph Neural Networks

Yixin Wu, Xinlei He, Pascal Berrang +4

A graph neural network (GNN) is a type of neural network that is specifically designed to process graph-structured data. Typically, GNNs can be implemented in two settings, includi…

cs.CR2025

SoK: Data Reconstruction Attacks Against Machine Learning Models: Definition, Metrics, and Benchmark

Rui Wen, Yiyong Liu, Michael Backes +1

Data reconstruction attacks, which aim to recover the training dataset of a target model with limited access, have gained increasing attention in recent years. However, there is cu…

cs.CR2023

Data Poisoning Attacks Against Multimodal Encoders

Ziqing Yang, Xinlei He, Zheng Li +4

Recently, the newly emerged multimodal models, which leverage both visual and linguistic modalities to train powerful encoders, have gained increasing attention. However, learning…

cs.CV2023

Generative Watermarking Against Unauthorized Subject-Driven Image Synthesis

Yihan Ma, Zhengyu Zhao, Xinlei He +3

Large text-to-image models have shown remarkable performance in synthesizing high-quality images. In particular, the subject-driven model makes it possible to personalize the image…

cs.CR2022

Dynamic Backdoor Attacks Against Machine Learning Models

Ahmed Salem, Rui Wen, Michael Backes +2

Machine learning (ML) has made tremendous progress during the past decade and is being adopted in various critical real-world applications. However, recent research has shown that…

cs.CR2023

SecurityNet: Assessing Machine Learning Vulnerabilities on Public Models

Boyang Zhang, Zheng Li, Ziqing Yang +4

While advanced machine learning (ML) models are deployed in numerous real-world applications, previous works demonstrate these models have security and privacy vulnerabilities. Var…

cs.CR2022

UnGANable: Defending Against GAN-based Face Manipulation

Zheng Li, Ning Yu, Ahmed Salem +3

Deepfakes pose severe threats of visual misinformation to our society. One representative deepfake application is face manipulation that modifies a victim's facial attributes in an…

cs.CR2021

BadNL: Backdoor Attacks against NLP Models with Semantic-preserving Improvements

Xiaoyi Chen, Ahmed Salem, Dingfan Chen +5

Deep neural networks (DNNs) have progressed rapidly during the past decade and have been deployed in various real-world applications. Meanwhile, DNN models have been shown to be vu…

cs.CR2024

Membership Inference Attacks Against In-Context Learning

Rui Wen, Zheng Li, Michael Backes +1

Adapting Large Language Models (LLMs) to specific tasks introduces concerns about computational efficiency, prompting an exploration of efficient methods such as In-Context Learnin…

cs.CR2021

ML-Doctor: Holistic Risk Assessment of Inference Attacks Against Machine Learning Models

Yugeng Liu, Rui Wen, Xinlei He +6

Inference attacks against Machine Learning (ML) models allow adversaries to learn sensitive information about training data, model parameters, etc. While researchers have studied,…

astro-ph.IM2021

How can astrotourism serve the sustainable development goals? The Namibian example

Hannah Dalgleish, Getachew Mengistie, Michael Backes +2

Astrotourism brings new opportunities to generate sustainable socio-economic development, preserve cultural heritage, and inspire and educate the citizens of the globe. This form o…

cs.CR2025

Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions

Yiting Qu, Ziqing Yang, Yihan Ma +3

Recent advances in text-to-image diffusion models have enabled the creation of a new form of digital art: optical illusions--visual tricks that create different perceptions of real…

cs.CR2020

Trollthrottle -- Raising the Cost of Astroturfing

Ilkan Esiyok, Lucjan Hanzlik, Robert Kuennemann +2

Astroturfing, i.e., the fabrication of public discourse by private or state-controlled sponsors via the creation of fake online accounts, has become incredibly widespread in recent…

cs.PL2025

Do You Even Lift? Strengthening Compiler Security Guarantees Against Spectre Attacks

Xaver Fabian, Marco Patrignani, Marco Guarnieri +1

Mainstream compilers implement different countermeasures to prevent specific classes of speculative execution attacks. Unfortunately, these countermeasures either lack formal guara…

cs.CR2023

Generated Graph Detection

Yihan Ma, Zhikun Zhang, Ning Yu +4

Graph generative models become increasingly effective for data distribution approximation and data augmentation. While they have aroused public concerns about their malicious misus…

cs.CR2020

BAAAN: Backdoor Attacks Against Autoencoder and GAN-Based Machine Learning Models

Ahmed Salem, Yannick Sautter, Michael Backes +2

The tremendous progress of autoencoders and generative adversarial networks (GANs) has led to their application to multiple critical tasks, such as fraud detection and sanitized da…

cs.CR2025

JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs

Junjie Chu, Yugeng Liu, Ziqing Yang +3

Jailbreak attacks aim to bypass the LLMs' safeguards. While researchers have proposed different jailbreak attacks in depth, they have done so in isolation -- either with unaligned…

cs.CR2025

Watermarking LLM-Generated Datasets in Downstream Tasks

Yugeng Liu, Tianshuo Cong, Michael Backes +2

Large Language Models (LLMs) have experienced rapid advancements, with applications spanning a wide range of fields, including sentiment classification, review generation, and ques…

cs.CR2024

"Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models

Xinyue Shen, Zeyuan Chen, Michael Backes +2

The misuse of large language models (LLMs) has drawn significant attention from the general public and LLM vendors. One particular type of adversarial prompt, known as jailbreak pr…

cs.SI2026

"Humans welcome to observe": A First Look at the Agent Social Network Moltbook

Yukun Jiang, Yage Zhang, Xinyue Shen +2

The rapid advancement of artificial intelligence (AI) agents has catalyzed the transition from static language models to autonomous agents capable of tool use, long-term planning,…

cs.CR2024

SOS! Soft Prompt Attack Against Open-Source Large Language Models

Ziqing Yang, Michael Backes, Yang Zhang +1

Open-source large language models (LLMs) have become increasingly popular among both the general public and industry, as they can be customized, fine-tuned, and freely used. Howeve…

cs.CY2023

Comprehensive Assessment of Toxicity in ChatGPT

Boyang Zhang, Xinyue Shen, Wai Man Si +6

Moderating offensive, hateful, and toxic language has always been an important but challenging topic in the domain of safe use in NLP. The emerging large language models (LLMs), su…

cs.LO2019

Causality & Control Flow

Robert Künnemann, Deepak Garg, Michael Backes

Causality has been the issue of philosophic debate since Hippocrates. It is used in formal verification and testing, e.g., to explain counterexamples or construct fault trees. Rece…

cs.CR2025

Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities

Yiting Qu, Michael Backes, Yang Zhang

Vision-language models (VLMs) are increasingly applied to identify unsafe or inappropriate images due to their internal ethical standards and powerful reasoning abilities. However,…

cs.CR2025

The Challenge of Identifying the Origin of Black-Box Large Language Models

Ziqing Yang, Yixin Wu, Yun Shen +3

The tremendous commercial potential of large language models (LLMs) has heightened concerns about their unauthorized use. Third parties can customize LLMs through fine-tuning and o…

cs.CR2026

Robustness Over Time: Understanding Adversarial Examples' Effectiveness on Longitudinal Versions of Large Language Models

Yugeng Liu, Tianshuo Cong, Zhengyu Zhao +3

Large Language Models (LLMs) undergo continuous updates to improve user experience. However, prior research on the security and safety implications of LLMs has primarily focused on…

cs.CR2020

On the security relevance of weights in deep learning

Kathrin Grosse, Thomas A. Trost, Marius Mosbach +2

Recently, a weight-based attack on stochastic gradient descent inducing overfitting has been proposed. We show that the threat is broader: A task-independent permutation on the ini…

cs.CR2015

PriCL: Creating a Precedent A Framework for Reasoning about Privacy Case Law

Michael Backes, Fabian Bendun, Joerg Hoffmann +1

We introduce PriCL: the first framework for expressing and automatically reasoning about privacy case law by means of precedent. PriCL is parametric in an underlying logic for expr…

cs.CR2025

PSGraph: Differentially Private Streaming Graph Synthesis by Considering Temporal Dynamics

Quan Yuan, Zhikun Zhang, Linkang Du +6

Streaming graphs are ubiquitous in daily life, such as evolving social networks and dynamic communication systems. Due to the sensitive information contained in the graph, directly…

cs.LG2024

Memorization in Self-Supervised Learning Improves Downstream Generalization

Wenhao Wang, Muhammad Ahmad Kaleem, Adam Dziedzic +3

Self-supervised learning (SSL) has recently received significant attention due to its ability to train high-performance encoders purely on unlabeled data-often scraped from the int…

cs.CR2026

Pop Quiz Attack: Black-box Membership Inference Attacks Against Large Language Models

Zeyuan Chen, Yihan Ma, Xinyue Shen +2

Large language models (LLMs) show strong performance across many applications, but their ability to memorize and potentially reveal training data raises serious privacy concerns. W…

astro-ph.IM2021

Dark sky tourism and sustainable development in Namibia

Hannah Dalgleish, Getachew Mengistie, Michael Backes +2

Namibia is world-renowned for its incredibly dark skies by the astronomy community, and yet, the country is not well recognised as a dark sky destination by tourists and travellers…

cs.CL2026

TrustLDM: Benchmarking Trustworthiness in Language Diffusion Models

Yichuan Mo, Yukun Jiang, Yanbo Shi +4

The rapid development of Language Diffusion Models (LDMs) challenges the dominant position of auto-regressive competitors in language processing. However, their flexible, any-order…

cs.CR2012

Adding Query Privacy to Robust DHTs

Michael Backes, Ian Goldberg, Aniket Kate +1

Interest in anonymous communication over distributed hash tables (DHTs) has increased in recent years. However, almost all known solutions solely aim at achieving sender or request…

cs.CR2024

Understanding Data Importance in Machine Learning Attacks: Does Valuable Data Pose Greater Harm?

Rui Wen, Michael Backes, Yang Zhang

Machine learning has revolutionized numerous domains, playing a crucial role in driving advancements and enabling data-centric processes. The significance of data in training model…

cs.CR2024

: Measuring Stereotypical Bias in Large Vision-Language Models from Vision and Language Modalities

Yukun Jiang, Zheng Li, Xinyue Shen +3

Large vision-language models (LVLMs) have been rapidly developed and widely used in various fields, but the (potential) stereotypical bias in the model is largely unexplored. In th…

cs.SI2020

Everything About You: A Multimodal Approach towards Friendship Inference in Online Social Networks

Tahleen Rahman, Mario Fritz, Michael Backes +1

Most previous works in privacy of Online Social Networks (OSN) focus on a restricted scenario of using one type of information to infer another type of information or using only st…

cs.CR2019

Towards Plausible Graph Anonymization

Yang Zhang, Mathias Humbert, Bartlomiej Surma +3

Social graphs derived from online social interactions contain a wealth of information that is nowadays extensively used by both industry and academia. However, as social graphs con…

cs.LG2024

Localizing Memorization in SSL Vision Encoders

Wenhao Wang, Adam Dziedzic, Michael Backes +1

Recent work on studying memorization in self-supervised learning (SSL) suggests that even though SSL encoders are trained on millions of images, they still memorize individual data…

cs.CR2024

Secure Composition of Robust and Optimising Compilers

Matthis Kruse, Michael Backes, Marco Patrignani

To ensure that secure applications do not leak their secrets, they are required to uphold several security properties such as spatial and temporal memory safety as well as cryptogr…

cs.CV2025

DivTrackee versus DynTracker: Promoting Diversity in Anti-Facial Recognition against Dynamic FR Strategy

Wenshu Fan, Minxing Zhang, Hongwei Li +5

The widespread adoption of facial recognition (FR) models raises serious concerns about their potential misuse, motivating the development of anti-facial recognition (AFR) to prote…

cs.CR2020

Killing four birds with one Gaussian process: the relation between different test-time attacks

Kathrin Grosse, Michael T. Smith, Michael Backes

In machine learning (ML) security, attacks like evasion, model stealing or membership inference are generally studied in individually. Previous work has also shown a relationship b…

astro-ph.IM2021

Astronomy outreach in Namibia: H.E.S.S. and beyond

Hannah Dalgleish, Heike Prokoph, Sylvia Zhu +6

Astronomy plays a major role in the scientific landscape of Namibia. Because of its excellent sky conditions, Namibia is home to ground-based observatories like the High Energy Spe…

cs.CR2014

Android Security Framework: Enabling Generic and Extensible Access Control on Android

Michael Backes, Sven Bugiel, Sebastian Gerling +1

We introduce the Android Security Framework (ASF), a generic, extensible security framework for Android that enables the development and integration of a wide spectrum of security…

cs.CR2025

HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns

Xinyue Shen, Yixin Wu, Yiting Qu +3

Large Language Models (LLMs) have raised increasing concerns about their misuse in generating hate speech. Among all the efforts to address this issue, hate speech detectors play a…

cs.LG2024

Open LLMs are Necessary for Current Private Adaptations and Outperform their Closed Alternatives

Vincent Hanke, Tom Blanchard, Franziska Boenisch +3

While open Large Language Models (LLMs) have made significant progress, they still fall short of matching the performance of their closed, proprietary counterparts, making the latt…

cs.CR2023

Mondrian: Prompt Abstraction Attack Against Large Language Models for Cheaper API Pricing

Wai Man Si, Michael Backes, Yang Zhang

The Machine Learning as a Service (MLaaS) market is rapidly expanding and becoming more mature. For example, OpenAI's ChatGPT is an advanced large language model (LLM) that generat…

cs.CL2026

PeerCheck: Enhancing LLM-Generated Academic Reviews Towards Human-Level Quality

Zeyuan Chen, Ziqing Yang, Yihan Ma +2

As academic submissions grow, the traditional peer review process struggles to keep up, raising concerns about quality and fairness. A trend of using large language models (LLMs) f…

cs.CR2016

Adversarial Perturbations Against Deep Neural Networks for Malware Classification

Kathrin Grosse, Nicolas Papernot, Praveen Manoharan +2

Deep neural networks, like many other machine learning models, have recently been shown to lack robustness against adversarially crafted inputs. These inputs are derived from regul…

cs.CR2023

Bilingual Problems: Studying the Security Risks Incurred by Native Extensions in Scripting Languages

Cristian-Alexandru Staicu, Sazzadur Rahaman, Ágnes Kiss +1

Scripting languages are continuously gaining popularity due to their ease of use and the flourishing software ecosystems that surround them. These languages offer crash and memory…

cs.CR2024

ICLGuard: Controlling In-Context Learning Behavior for Applicability Authorization

Wai Man Si, Michael Backes, Yang Zhang

In-context learning (ICL) is a recent advancement in the capabilities of large language models (LLMs). This feature allows users to perform a new task without updating the model. C…

astro-ph.IM2019

Millimeter-wave Monitoring of Active Galactic Nuclei with the Africa Millimetre Telescope

Michael Backes, Markus Böttcher, Heino Falcke

Active Galactic Nuclei are the dominant sources of gamma rays outside our Galaxy and also candidates for being the source of ultra-high energy cosmic rays. In addition to being emi…

cs.CR2026

BadBone: Backdoor Attacks Against Backbone Models in Visual Prompt Learning

Ziqing Yang, Rui Wen, Xinlei He +3

Prompt learning is a new machine learning paradigm that has attracted ample attention due to its simplicity and proven efficacy. Despite its growing adoption, the security vulnerab…

cs.CR2025

UnsafeBench: Benchmarking Image Safety Classifiers on Real-World and AI-Generated Images

Yiting Qu, Xinyue Shen, Yixin Wu +3

With the advent of text-to-image models and concerns about their misuse, developers are increasingly relying on image safety classifiers to moderate their generated unsafe images.…

cs.CR2025

Amplifying Machine Learning Attacks Through Strategic Compositions

Yugeng Liu, Zheng Li, Hai Huang +2

Machine learning (ML) models are proving to be vulnerable to a variety of attacks that allow the adversary to learn sensitive information, cause mispredictions, and more. While the…

cs.CR2022

Mental Models of Adversarial Machine Learning

Lukas Bieringer, Kathrin Grosse, Michael Backes +2

Although machine learning is widely used in practice, little is known about practitioners' understanding of potential security challenges. In this work, we close this substantial g…

cs.CR2024

Prompt Stealing Attacks Against Text-to-Image Generation Models

Xinyue Shen, Yiting Qu, Michael Backes +1

Text-to-Image generation models have revolutionized the artwork design process and enabled anyone to create high-quality images by entering text descriptions called prompts. Creati…

cs.CR2016

A Survey on Routing in Anonymous Communication Protocols

Fatemeh Shirazi, Milivoj Simeonovski, Muhammad Rizwan Asghar +2

The Internet has undergone dramatic changes in the past 15 years, and now forms a global communication platform that billions of users rely on for their daily activities. While thi…

cs.CR2021

Towards a Principled Approach for Dynamic Analysis of Android's Middleware

Oliver Schranz, Sebastian Weisgerber, Erik Derr +2

The Android middleware, in particular the so-called systemserver, is a crucial and central component to Android's security and robustness. To understand whether the systemserver pr…

cs.CR2021

When Machine Unlearning Jeopardizes Privacy

Min Chen, Zhikun Zhang, Tianhao Wang +3

The right to be forgotten states that a data owner has the right to erase their data from an entity storing it. In the context of machine learning (ML), the right to be forgotten r…

cs.CR2011

X-pire! - A digital expiration date for images in social networks

Julian Backes, Michael Backes, Markus Dürmuth +2

The Internet and its current information culture of preserving all kinds of data cause severe problems with privacy. Most of today's Internet users, especially teenagers, publish v…

cs.CR2025

GEO-Detective: Unveiling Location Privacy Risks in Images with LLM Agents

Xinyu Zhang, Yixin Wu, Boyang Zhang +4

Images shared on social media often expose geographic cues. While early geolocation methods required expert effort and lacked generalization, the rise of Large Vision Language Mode…

cs.CR2023

In ChatGPT We Trust? Measuring and Characterizing the Reliability of ChatGPT

Xinyue Shen, Zeyuan Chen, Michael Backes +1

The way users acquire information is undergoing a paradigm shift with the advent of ChatGPT. Unlike conventional search engines, ChatGPT retrieves knowledge from the model itself a…

astro-ph.HE2011

Monitoring of bright, nearby Active Galactic Nuclei with the MAGIC telescopes

Robert Wagner, Michael Backes, Konstancja Satalecka +7

Observations and detections of Active Galactic Nuclei (AGN) by Cherenkov telescopes are often triggered by information about high flux states in other wavelength bands. To overcome…

cs.CR2025

Revisiting Transferable Adversarial Images: Systemization, Evaluation, and New Insights

Zhengyu Zhao, Hanwei Zhang, Renjue Li +6

Transferable adversarial images raise critical security concerns for computer vision systems in real-world, black-box attack scenarios. Although many transfer attacks have been pro…