activity
20242026
collaborators

5 papers

cs.CV2026

LaV-CoT: Language-Aware Visual CoT with Multi-Aspect Reward Optimization for Real-World Multilingual VQA

Jing Huang, Zhiya Tan, Shutao Gong +6

As large vision language models (VLMs) advance, their capabilities in multilingual visual question answering (mVQA) have significantly improved. Chain-of-thought (CoT) reasoning ha…

cs.CV2026

DocShield: Towards AI Document Safety via Evidence-Grounded Agentic Reasoning

Fanwei Zeng, Changtao Miao, Jing Huang +9

The rapid progress of generative AI has enabled increasingly realistic text-centric image forgeries, posing major challenges to document safety. Existing forensic methods mainly re…

cs.AI2025

LogicLens: Visual-Logical Co-Reasoning for Text-Centric Forgery Analysis

Fanwei Zeng, Changtao Miao, Jing Huang +7

Sophisticated text-centric forgeries, fueled by rapid AIGC advancements, pose a significant threat to societal security and information authenticity. Current methods for text-centr…

cs.CV2025

DDL: A Large-Scale Datasets for Deepfake Detection and Localization in Diversified Real-World Scenarios

Changtao Miao, Yi Zhang, Weize Gao +11

Recent advances in AIGC have exacerbated the misuse of malicious deepfake content, making the development of reliable deepfake detection methods an essential means to address this…

cs.CV2025

Modelship Attribution: Tracing Multi-Stage Manipulations Across Generative Models

Zhiya Tan, Xin Zhang, Joey Tianyi Zhou

As generative techniques become increasingly accessible, authentic visuals are frequently subjected to iterative alterations by various individuals employing a variety of tools. Cu…