activity
20242026
collaborators

8 papers

cs.CV2026

ForensicsTok: Forensics-Guided Tokenized Modeling for Image Tampering Localization

Lei Xu, Haowei Wang, Shen Chen +3

Multi-modal Large Language Models (MLLMs) offer powerful reasoning for forensic tasks, yet existing approaches utilizing exogenous segmentation decoders often suffer from suboptima…

cs.CR2026

When Backdoors Go Beyond Triggers: Semantic Drift in Diffusion Models Under Encoder Attacks

Shenyang Chen, Liuwan Zhu

Standard evaluations of backdoor attacks on text-to-image (T2I) models primarily measure trigger activation and visual fidelity. We challenge this paradigm, demonstrating that enco…

cs.CV2026

ForgeryVCR: Visual-Centric Reasoning via Efficient Forensic Tools in MLLMs for Image Forgery Detection and Localization

Youqi Wang, Shen Chen, Haowei Wang +6

Existing Multimodal Large Language Models (MLLMs) for image forgery detection and localization predominantly operate under a text-centric Chain-of-Thought (CoT) paradigm. However,…

cs.CV2025

TripleFDS: Triple Feature Disentanglement and Synthesis for Scene Text Editing

Yuchen Bao, Yiting Wang, Wenjian Huang +5

Scene Text Editing (STE) aims to naturally modify text in images while preserving visual consistency, the decisive factors of which can be divided into three parts, i.e., text styl…

cs.CV2025

Orthogonal Subspace Decomposition for Generalizable AI-Generated Image Detection

Zhiyuan Yan, Jiangming Wang, Peng Jin +7

AI-generated images (AIGIs), such as natural or face images, have become increasingly important yet challenging. In this paper, we start from a new perspective to excavate the reas…

cs.CV2025

A Quality-Centric Framework for Generic Deepfake Detection

Wentang Song, Zhiyuan Yan, Yuzhen Lin +6

Detecting AI-generated images, particularly deepfakes, has become increasingly crucial, with the primary challenge being the generalization to previously unseen manipulation method…