From the 1 of 20 linked papers with an AI index.
20 papers
The Boy Who Cried Wolf: Adversarial Misclassification of Safe Inputs as Unsafe in Multimodal Guardrails
Shuo Shi, Rui Yin, Naen Xu +7
Multimodal guard models have emerged as critical safety components for screening content in vision-language systems. While adversarial research has extensively studied jailbreaking…
Lilith: Backdoor Generalization under Training-Inference Trigger Shift
Zhou Feng, Jiahao Chen, Chunyi Zhou +6
The paper studies how backdoor attacks can remain effective when the trigger used at inference time differs from the one seen during training, and proposes Lilith, a black‑box meth…
Decision-Level Hijacking: Injecting Cognitive Bias into Large Language Models via Bit-Flip Attacks
Yu Yan, Jiahao Chen, Siqi Lu +6
Large Language Models (LLMs) have been widely applied in high-stakes decision-making scenarios such as corporate strategy, and users are increasingly relying on their outputs. Howe…
LoRAShield: Data-Free Editing Alignment for Secure Personalized LoRA Sharing
Jiahao Chen, Junhao Li, Yiming Wang +6
The proliferation of Low-Rank Adaptation (LoRA) models has democratized personalized text-to-image generation, enabling users to share lightweight models (e.g., personal portraits)…
Customization under Fire: Plugin Poisoning in Text-to-Image Ecosystem
Jiahao Chen, Xing He, Yong Yang +6
The prosperity of text-to-image (T2I) models has fostered a vibrant share-and-play ecosystem centered on Low-Rank Adaptation (LoRA) plugins, which allow users to customize and shar…
Order within Chaos: Capturing Intrinsic Energy Anomalies for AI-Manipulated Image Forgery Localization
Yiming Wang, Baiqi Wu, Qingming Li +3
Recent advancements in generative AI have led to image editing models capable of producing realistic forgeries that evade traditional image forgery localization methods, as these a…