2 papers
cs.AI2026
Every Picture Tells a Dangerous Story: Memory-Augmented Multi-Agent Jailbreak Attacks on VLMs
Jianhao Chen, Haoyang Chen, Shiqin Wang +4
Vision-Language Models (VLMs) expand the attack surface of safety-aligned systems by coupling visual perception with text generation. Existing multimodal jailbreak attacks primaril…
cs.CV2025
SEdit: Text-Guided Image Editing with Precise Semantic and Spatial Control
Xudong Liu, Zikun Chen, Ruowei Jiang +5
Recent advances in diffusion models have enabled high-quality generation and manipulation of images guided by texts, as well as concept learning from images. However, naive applica…