7 papers
When the Prompt Becomes Visual: Vision-Centric Jailbreak Attacks for Large Image Editing Models
Jiacheng Hou, Yining Sun, Ruochong Jin +4
Recent advances in large image editing models have shifted the paradigm from text-driven instructions to vision-prompt editing, where user intent is inferred directly from visual i…
VPA-Guard: Defending and Benchmarking Image-to-Video Generation Against Visual Prompt Attacks
Yining Sun, Haoyu Kang, Jiajun Wu +7
Recent advancements in Image-to-Video (I2V) generation have transformed input images from simple appearance references into interactive control interfaces where visual cues such as…
Evaluating and Enhancing Negation Comprehension in Remote Sensing MLLMs
Haochen Han, Jue Wang, Alex Jinpeng Wang +1
Multimodal Large Language Models (MLLMs) have demonstrated remarkable success in various Remote Sensing (RS) tasks. However, their ability to comprehend negation remains underexplo…
TextEditBench: Evaluating Reasoning-aware Text Editing Beyond Rendering
Rui Gui, Yang Wan, Haochen Han +4
Text rendering has recently emerged as one of the most challenging frontiers in visual generation, drawing significant attention from large-scale diffusion and multimodal models. H…
Negation-Aware Test-Time Adaptation for Vision-Language Models
Haochen Han, Alex Jinpeng Wang, Fangming Liu +1
In this paper, we study a practical but less-touched problem in Vision-Language Models (VLMs), \ie, negation understanding. Specifically, many real-world applications require model…
Unlearning the Noisy Correspondence Makes CLIP More Robust
Haochen Han, Alex Jinpeng Wang, Peijun Ye +1
The data appetite for Vision-Language Models (VLMs) has continuously scaled up from the early millions to billions today, which faces an untenable trade-off with data quality and i…