collaborators

8 papers

cs.CR2026

How China-Origin Vision-Language Models Move from Refusal to Reframing in State Alignment

Guang Yang, Fengchen Liu, Alex Wang +2

State-aligned distortion has been documented in China-origin text-based large language models (LLMs), but whether, and in what form, it arises in multimodal systems has not been sy…

cs.CV2026

When the Prompt Becomes Visual: Vision-Centric Jailbreak Attacks for Large Image Editing Models

Jiacheng Hou, Yining Sun, Ruochong Jin +4

Recent advances in large image editing models have shifted the paradigm from text-driven instructions to vision-prompt editing, where user intent is inferred directly from visual i…

cs.CV2026

VPA-Guard: Defending and Benchmarking Image-to-Video Generation Against Visual Prompt Attacks

Yining Sun, Haoyu Kang, Jiajun Wu +7

Recent advancements in Image-to-Video (I2V) generation have transformed input images from simple appearance references into interactive control interfaces where visual cues such as…

cs.CV2026

Evaluating and Enhancing Negation Comprehension in Remote Sensing MLLMs

Haochen Han, Jue Wang, Alex Jinpeng Wang +1

Multimodal Large Language Models (MLLMs) have demonstrated remarkable success in various Remote Sensing (RS) tasks. However, their ability to comprehend negation remains underexplo…

cs.CV2025

TextEditBench: Evaluating Reasoning-aware Text Editing Beyond Rendering

Rui Gui, Yang Wan, Haochen Han +4

Text rendering has recently emerged as one of the most challenging frontiers in visual generation, drawing significant attention from large-scale diffusion and multimodal models. H…

cs.CV2025

Skeletons Speak Louder than Text: A Motion-Aware Pretraining Paradigm for Video-Based Person Re-Identification

Rifen Lin, Alex Jinpeng Wang, Jiawei Mo +1

Multimodal pretraining has revolutionized visual understanding, but its impact on video-based person re-identification (ReID) remains underexplored. Existing approaches often rely…