7 papers
GUIGuard-Bench: Toward a General Evaluation for Privacy-Preserving GUI Agents
Yanxi Wang, Zhiling Zhang, Wenbo Zhou +6
As GUI agents increasingly rely on screenshots to perceive and operate digital environments, they may inadvertently expose sensitive information such as identities, accounts, locat…
VertMark: A Unified Training-Free Robust Watermarking Framework for Vertical Domain Pre-trained Language Models
Cong Kong, Xin Cheng, Zhaoxia Yin +3
With the application of vertical domain pre-trained language models (VPLMs) in specialized fields such as medical, finance, and law, model parameters and inference capabilities hav…
State-Dependent Safety Failures in Multi-Turn Language Model Interaction
Pengcheng Li, Jie Zhang, Tianwei Zhang +5
Safety alignment in large language models is typically evaluated under isolated queries, yet real-world use is inherently multi-turn. Although multi-turn jailbreaks are empirically…
Character as a Latent Variable in Large Language Models: A Mechanistic Account of Emergent Misalignment and Conditional Safety Failures
Yanghao Su, Wenbo Zhou, Tianwei Zhang +4
Emergent Misalignment refers to a failure mode in which fine-tuning large language models (LLMs) on narrowly scoped data induces broadly misaligned behavior. Prior explanations mai…
SafeGuider: Robust and Practical Content Safety Control for Text-to-Image Models
Peigui Qi, Kunsheng Tang, Wenbo Zhou +5
Text-to-image models have shown remarkable capabilities in generating high-quality images from natural language descriptions. However, these models are highly vulnerable to adversa…
BURN: Backdoor Unlearning via Adversarial Boundary Analysis
Yanghao Su, Jie Zhang, Yiming Li +6
Backdoor unlearning aims to remove backdoor-related information while preserving the model's original functionality. However, existing unlearning methods mainly focus on recovering…