8 papers
ForensicsTok: Forensics-Guided Tokenized Modeling for Image Tampering Localization
Lei Xu, Haowei Wang, Shen Chen +3
Multi-modal Large Language Models (MLLMs) offer powerful reasoning for forensic tasks, yet existing approaches utilizing exogenous segmentation decoders often suffer from suboptima…
When Backdoors Go Beyond Triggers: Semantic Drift in Diffusion Models Under Encoder Attacks
Shenyang Chen, Liuwan Zhu
Standard evaluations of backdoor attacks on text-to-image (T2I) models primarily measure trigger activation and visual fidelity. We challenge this paradigm, demonstrating that enco…
ForgeryVCR: Visual-Centric Reasoning via Efficient Forensic Tools in MLLMs for Image Forgery Detection and Localization
Youqi Wang, Shen Chen, Haowei Wang +6
Existing Multimodal Large Language Models (MLLMs) for image forgery detection and localization predominantly operate under a text-centric Chain-of-Thought (CoT) paradigm. However,…
TripleFDS: Triple Feature Disentanglement and Synthesis for Scene Text Editing
Yuchen Bao, Yiting Wang, Wenjian Huang +5
Scene Text Editing (STE) aims to naturally modify text in images while preserving visual consistency, the decisive factors of which can be divided into three parts, i.e., text styl…
Orthogonal Subspace Decomposition for Generalizable AI-Generated Image Detection
Zhiyuan Yan, Jiangming Wang, Peng Jin +7
AI-generated images (AIGIs), such as natural or face images, have become increasingly important yet challenging. In this paper, we start from a new perspective to excavate the reas…
A Quality-Centric Framework for Generic Deepfake Detection
Wentang Song, Zhiyuan Yan, Yuzhen Lin +6
Detecting AI-generated images, particularly deepfakes, has become increasingly crucial, with the primary challenge being the generalization to previously unseen manipulation method…