5 papers
LogicLens: Visual-Logical Co-Reasoning for Text-Centric Forgery Analysis
Fanwei Zeng, Changtao Miao, Jing Huang +7
Sophisticated text-centric forgeries, fueled by rapid AIGC advancements, pose a significant threat to societal security and information authenticity. Current methods for text-centr…
LaV-CoT: Language-Aware Visual CoT with Multi-Aspect Reward Optimization for Real-World Multilingual VQA
Jing Huang, Zhiya Tan, Shutao Gong +6
As large vision language models (VLMs) advance, their capabilities in multilingual visual question answering (mVQA) have significantly improved. Chain-of-thought (CoT) reasoning ha…
Modelship Attribution: Tracing Multi-Stage Manipulations Across Generative Models
Zhiya Tan, Xin Zhang, Joey Tianyi Zhou
As generative techniques become increasingly accessible, authentic visuals are frequently subjected to iterative alterations by various individuals employing a variety of tools. Cu…
DDL: A Large-Scale Datasets for Deepfake Detection and Localization in Diversified Real-World Scenarios
Changtao Miao, Yi Zhang, Weize Gao +11
Recent advances in AIGC have exacerbated the misuse of malicious deepfake content, making the development of reliable deepfake detection methods an essential means to address this…
Generative adversarial learning with optimal input dimension and its adaptive generator architecture
Zhiyao Tan, Ling Zhou, Huazhen Lin
We investigate the impact of the input dimension on the generalization error in generative adversarial networks (GANs). In particular, we first provide both theoretical and practic…