5 papers
REVEAL: Reference-Grounded Reasoning for Multimodal Manipulation Detection
Jun Zhou, Bingwen Hu, Yaxiong Wang +4
Multimodal manipulation detection aims to simultaneously identify forged image--text pairs and localize tampered regions, yet existing methods typically rely on memorizing isolated…
OmniVL-Guard Pro: A Tool-Augmented Agent for Omnibus Vision-Language Forensics
Jinjie Shen, Zheng Huang, Yuchen Zhang +7
Existing vision-language forgery detection and grounding methods operate under a closed-world paradigm, assuming verification can be completed by the model alone. However, self-con…
MultiHateLoc: Towards Temporal Localisation of Multimodal Hate Content in Online Videos
Qiyue Sun, Tailin Chen, Yinghui Zhang +4
The rapid growth of video content on platforms such as TikTok and YouTube has intensified the spread of multimodal hate speech, where harmful cues emerge subtly and asynchronously…
Training-Free and Interpretable Hateful Video Detection via Multi-stage Adversarial Reasoning
Shuonan Yang, Yuchen Zhang, Zeyu Fu
Hateful videos pose serious risks by amplifying discrimination, inciting violence, and undermining online safety. Existing training-based hateful video detection methods are constr…
Enhanced Multimodal Hate Video Detection via Channel-wise and Modality-wise Fusion
Yinghui Zhang, Tailin Chen, Yuchen Zhang +1
The rapid rise of video content on platforms such as TikTok and YouTube has transformed information dissemination, but it has also facilitated the spread of harmful content, partic…