activity
20242026
most citedDeepfake Detection that Generalizes Across Benchmarks

7 citations · 8 across the 2 of their papers we have counts for

collaborators

9 papers

cs.CV20267 cited

Deepfake Detection that Generalizes Across Benchmarks

Andrii Yermakov, Jan Cech, Jiri Matas +1

The generalization of deepfake detectors to unseen manipulation techniques remains a challenge for practical deployment. Although many approaches adapt foundation models by introdu…

cs.LG20261 cited

Red-Teaming Text-to-Image Models via In-Context Experience Replay and Semantic-Preserving Prompt Rewriting

Zhi-Yi Chin, Pin-Yu Chen, Wei-Chen Chiu +1

Understanding the capabilities of text-to-image (T2I) models in harmful content generation is essential to safety and compliance. However, human red-teaming is costly and inconsist…

cs.CR2026

ProxyPrompt: Securing System Prompts against Prompt Extraction Attacks

Zhixiong Zhuang, Maria-Irina Nicolae, Hui-Po Wang +1

The integration of large language models (LLMs) into a wide range of applications has highlighted the critical role of well-crafted system prompts, which require extensive testing…

cs.LG2025

Probe-based Fine-tuning for Reducing Toxicity

Jan Wehner, Mario Fritz

Probes trained on model activations can detect undesirable behaviors like deception or biases that are difficult to identify from outputs alone. This makes them useful detectors to…

cs.LG2025

DP-SNP-TIHMM: Differentially Private, Time-Inhomogeneous Hidden Markov Models for Synthesizing Genome-Wide Association Datasets

Shadi Rahimian, Mario Fritz

Single nucleotide polymorphism (SNP) datasets are fundamental to genetic studies but pose significant privacy risks when shared. The correlation of SNPs with each other makes stron…

cs.CR2025

Stealix: Model Stealing via Prompt Evolution

Zhixiong Zhuang, Hui-Po Wang, Maria-Irina Nicolae +1

Model stealing poses a significant security risk in machine learning by enabling attackers to replicate a black-box model without access to its training data, thus jeopardizing int…