3 papers
cs.CV2026
ForensicsTok: Forensics-Guided Tokenized Modeling for Image Tampering Localization
Lei Xu, Haowei Wang, Shen Chen +3
Multi-modal Large Language Models (MLLMs) offer powerful reasoning for forensic tasks, yet existing approaches utilizing exogenous segmentation decoders often suffer from suboptima…
cs.CV2026
ForgeryVCR: Visual-Centric Reasoning via Efficient Forensic Tools in MLLMs for Image Forgery Detection and Localization
Youqi Wang, Shen Chen, Haowei Wang +6
Existing Multimodal Large Language Models (MLLMs) for image forgery detection and localization predominantly operate under a text-centric Chain-of-Thought (CoT) paradigm. However,…
cs.CV2025
TripleFDS: Triple Feature Disentanglement and Synthesis for Scene Text Editing
Yuchen Bao, Yiting Wang, Wenjian Huang +5
Scene Text Editing (STE) aims to naturally modify text in images while preserving visual consistency, the decisive factors of which can be divided into three parts, i.e., text styl…