works on

From the 1 of 25 linked papers with an AI index.

activity
20242026
collaborators

25 papers

cs.CR2026

The Boy Who Cried Wolf: Adversarial Misclassification of Safe Inputs as Unsafe in Multimodal Guardrails

Shuo Shi, Rui Yin, Naen Xu +7

Multimodal guard models have emerged as critical safety components for screening content in vision-language systems. While adversarial research has extensively studied jailbreaking…

cs.CR2026

Lilith: Backdoor Generalization under Training-Inference Trigger Shift

Zhou Feng, Jiahao Chen, Chunyi Zhou +6

The paper studies how backdoor attacks can remain effective when the trigger used at inference time differs from the one seen during training, and proposes Lilith, a black‑box meth…

cs.CR2026

LoRAShield: Data-Free Editing Alignment for Secure Personalized LoRA Sharing

Jiahao Chen, Junhao Li, Yiming Wang +6

The proliferation of Low-Rank Adaptation (LoRA) models has democratized personalized text-to-image generation, enabling users to share lightweight models (e.g., personal portraits)…

cs.CR2026

Customization under Fire: Plugin Poisoning in Text-to-Image Ecosystem

Jiahao Chen, Xing He, Yong Yang +6

The prosperity of text-to-image (T2I) models has fostered a vibrant share-and-play ecosystem centered on Low-Rank Adaptation (LoRA) plugins, which allow users to customize and shar…

cs.LG2026

Angel or Demon: Investigating the Plasticity Interventions' Impact on Backdoor Threats in Deep Reinforcement Learning

Oubo Ma, Ruixiao Lin, Yang Dai +4

Extensive research has highlighted the severe threats posed by backdoor attacks to deep reinforcement learning (DRL). However, prior studies primarily focus on vanilla scenarios, w…

cs.CR2026

Profiling for Pennies: Unveiling the Privacy Iceberg of LLM Agents

Jiahao Chen, Qi Zhang, Ruixiao Lin +7

Large Language Models (LLMs) have revolutionized how information are collected, aggregated, and reasoned. However, this enables a novel and accessible vector of privacy intrusion:…