5 papers
State Backdoor: Towards Stealthy Real-world Poisoning Attack on Vision-Language-Action Model in State Space
Ji Guo, Wenbo Jiang, Yansong Lin +6
Vision-Language-Action (VLA) models are widely deployed in safety-critical embodied AI applications such as robotics. However, their complex multimodal interactions also expose new…
DIVER:Diving Deeper into Distilled Data via Expressive Semantic Recovery
Qianxin Xia, Zhiyong Shu, Wenbo Jiang +3
Dataset distillation aims to synthesize a compact proxy dataset that is unreadable or non-raw from the original dataset for privacy protection and highly efficient learning. Howeve…
AlignVAR: Towards Globally Consistent Visual Autoregression for Image Super-Resolution
Cencen Liu, Dongyang Zhang, Wen Yin +6
Visual autoregressive (VAR) models have recently emerged as a promising alternative for image generation, offering stable training, non-iterative inference, and high-fidelity synth…
TrojanEdit: Multimodal Backdoor Attack Against Image Editing Model
Ji Guo, Peihong Chen, Wenbo Jiang +6
Multimodal diffusion models for image editing generate outputs conditioned on both textual instructions and visual inputs, aiming to modify target regions while preserving the rest…
One Prompt to Verify Your Models: Black-Box Text-to-Image Models Verification via Non-Transferable Adversarial Attacks
Ji Guo, Wenbo Jiang, Rui Zhang +2
Recently, various types of Text-to-Image (T2I) models have emerged (such as DALL-E and Stable Diffusion), and showing their advantages in different aspects. Therefore, some third-p…