8 papers
StepGuard: Guarding Web Navigation via Single-Step Calibration
Zhihao Cui, Yuchen Zhang, Xiyang Sun +6
Web navigation requires agents to follow natural language goals, interact with web pages, and produce accurate answers. While recent advances leverage vision-language models and re…
CORE: Conflict-Oriented Reasoning for General Multimodal Manipulation Detection
Jinjie Shen, Yaxiong Wang, Yujiao Wu +5
The rapid rise of generative AI has made multimodal fake news increasingly realistic and pervasive, posing severe threats to public trust and social stability. Existing detection m…
OmniVL-Guard Pro: A Tool-Augmented Agent for Omnibus Vision-Language Forensics
Jinjie Shen, Zheng Huang, Yuchen Zhang +7
Existing vision-language forgery detection and grounding methods operate under a closed-world paradigm, assuming verification can be completed by the model alone. However, self-con…
Cultivating Forensic Reasoning for Generalizable Multimodal Manipulation Detection
Yuchen Zhang, Yaxiong Wang, Kecheng Han +4
Recent advances in generative AI have significantly enhanced the realism of multimodal media manipulation, thereby posing substantial challenges to manipulation detection. Existing…
Generating Attribution Reports for Manipulated Facial Images: A Dataset and Baseline
Jingchun Lian, Lingyu Liu, Yaxiong Wang +4
Existing facial forgery detection methods typically focus on binary classification or pixel-level localization, providing little semantic insight into the nature of the manipulatio…
Efficient Encoder-Free Fourier-based 3D Large Multimodal Model
Guofeng Mei, Wei Lin, Luigi Riz +3
Large Multimodal Models (LMMs) that process 3D data typically rely on heavy, pre-trained visual encoders to extract geometric features. While recent 2D LMMs have begun to eliminate…