collaborators

8 papers

cs.AI2026

StepGuard: Guarding Web Navigation via Single-Step Calibration

Zhihao Cui, Yuchen Zhang, Xiyang Sun +6

Web navigation requires agents to follow natural language goals, interact with web pages, and produce accurate answers. While recent advances leverage vision-language models and re…

cs.AI2026

CORE: Conflict-Oriented Reasoning for General Multimodal Manipulation Detection

Jinjie Shen, Yaxiong Wang, Yujiao Wu +5

The rapid rise of generative AI has made multimodal fake news increasingly realistic and pervasive, posing severe threats to public trust and social stability. Existing detection m…

cs.CV2026

OmniVL-Guard Pro: A Tool-Augmented Agent for Omnibus Vision-Language Forensics

Jinjie Shen, Zheng Huang, Yuchen Zhang +7

Existing vision-language forgery detection and grounding methods operate under a closed-world paradigm, assuming verification can be completed by the model alone. However, self-con…

cs.CV2026

Cultivating Forensic Reasoning for Generalizable Multimodal Manipulation Detection

Yuchen Zhang, Yaxiong Wang, Kecheng Han +4

Recent advances in generative AI have significantly enhanced the realism of multimodal media manipulation, thereby posing substantial challenges to manipulation detection. Existing…

cs.CV2026

Generating Attribution Reports for Manipulated Facial Images: A Dataset and Baseline

Jingchun Lian, Lingyu Liu, Yaxiong Wang +4

Existing facial forgery detection methods typically focus on binary classification or pixel-level localization, providing little semantic insight into the nature of the manipulatio…

cs.CV2026

Efficient Encoder-Free Fourier-based 3D Large Multimodal Model

Guofeng Mei, Wei Lin, Luigi Riz +3

Large Multimodal Models (LMMs) that process 3D data typically rely on heavy, pre-trained visual encoders to extract geometric features. While recent 2D LMMs have begun to eliminate…