6 papers
Open Models, Open Risks: Measuring Unsafe Generation in Text-to-Image Models In the Wild
Peilin Han, Yang Liu, Yilong Yang +4
Existing safety studies on text-to-image (T2I) jailbreaks are largely conducted in controlled in-the-lab settings, typically on a small number of canonical models. As a result, the…
Benign Inputs, Harmful Outputs: Cross-Modal Jailbreaking via Distributed Semantic Recomposition
Yani Wang, Yilong Yang, Yang Liu +3
Multimodal Large Language Models (MLLMs) have recently demonstrated remarkable capabilities in content synthesis and autonomous reasoning. Previous safety guardrails are primarily…
A Cross-Modal Prompt Injection Attack against Large Vision-Language Models with Image-Only Perturbation
Hao Yang, Zhuo Ma, Yang Liu +3
Large vision-language models (LVLMs) have emerged as a powerful paradigm for multimodal intelligence, but their growing deployment also expands the attack surface of prompt injecti…
On the Generation and Mitigation of Harmful Geometry in Image-to-3D Models
Yule Liu, Yilong Yang, Jiale Teng +11
Recent advances in image-to-3D models have significantly improved the fidelity and accessibility of 3D content creation. Such a powerful reconstruction capability that enables crea…
Auditing Data Membership in Reinforcement Learning With Verifiable Rewards
Yule Liu, Heyi Zhang, Jinyi Zheng +6
Reinforcement Learning with Verifiable Rewards (RLVR) has become a core training stage in recent large language models (LLMs). Its reliance on non-public, high-value prompt sets ra…
Improving Sustainability of Adversarial Examples in Class-Incremental Learning
Taifeng Liu, Xinjing Liu, Liangqiu Dong +3
Current adversarial examples (AEs) are typically designed for static models. However, with the wide application of Class-Incremental Learning (CIL), models are no longer static and…