40 papers
CHILLGuard: Towards Fine-Grained Chinese LLM Safety Guardrail with Scalable Data Construction and Model-aware Preference Alignment
Wenbo Yu, Bohua Wang, Hao Fang +10
Malicious content generated from large language models (LLMs) could pose severe safety risks and ethical concerns. While existing LLM safety guardrails excel in English or multilin…
MemVenom: Triggered Poisoning of Multimodal Memories in Web Agents
Yv Zhang, Hao Sun, Hao Fang +5
External memory has become a core component of modern web agents, enabling long-horizon reasoning through the retrieval of past experiences. However, this paradigm introduces a cri…
Bypassing Copyright Protection in Diffusion-based Customization via Two-Stage Latent Feature Optimization
Ziang Xu, Wenbo Yu, Hongyao Yu +6
With the growing concerns over copyright infringement in diffusion-based customization, adversarial attacks have emerged as a prominent defense strategy to prevent malicious conten…
Reasoning Matters: Mitigate Hallucination in Multimodal Large Reasoning Models via Reasoning-Conditioned Preference Optimization
Jiawei Kong, Hao Fang, Shunxiang Liao +5
Multimodal Large Reasoning Models introduce the reasoning paradigm, demonstrating strong capabilities on complex vision-language tasks. However, they still suffer from severe hallu…
FlowErase-RL: Rethinking Concept Erasure as Reward Optimization in Flow Matching Models
Yi Sun, Zhiqi Zhang, Xinhao Zhong +5
Recent advances in flow matching models have significantly improved text-to-image generation quality, but also introduce growing safety risks due to the generation of harmful or un…
CPC-VAR:Continual Personalized and Compositional Generation in Visual Autoregressive Models
Junhao Li, Xinhao Zhong, Yi sun +4
Visual autoregressive (VAR) models have recently emerged as an efficient paradigm for text-to-image generation. Despite their strong generative capability, existing VAR-based perso…