16 citations · 20 across the 29 of their papers we have counts for
14 papers · 1 filter
UniHetero: Could Generation Enhance Understanding for Vision-Language-Model at Large Data Scale?
Fengjiao Chen, Minhao Jing, Weitao Lu +3
Vision-language large models are moving toward the unification of visual understanding and visual generation tasks. However, whether generation can enhance understanding is still u…
Fine-grained Multiple Supervisory Network for Multi-modal Manipulation Detecting and Grounding
Xinquan Yu, Wei Lu, Xiangyang Luo
The task of Detecting and Grounding Multi-Modal Media Manipulation (DGM) is a branch of misinformation detection. Unlike traditional binary classification, it includes complex…
Weakly Supervised Multimodal Temporal Forgery Localization via Multitask Learning
Wenbo Xu, Wei Lu, Xiangyang Luo
The spread of Deepfake videos has caused a trust crisis and impaired social stability. Although numerous approaches have been proposed to address the challenges of Deepfake detecti…
A Multimodal Deviation Perceiving Framework for Weakly-Supervised Temporal Forgery Localization
Wenbo Xu, Junyan Wu, Wei Lu +2
Current researches on Deepfake forensics often treat detection as a classification task or temporal forgery localization problem, which are usually restrictive, time-consuming, and…
Weakly-Supervised Image Forgery Localization via Vision-Language Collaborative Reasoning Framework
Ziqi Sheng, Junyan Wu, Wei Lu +1
Image forgery localization aims to precisely identify tampered regions within images, but it commonly depends on costly pixel-level annotations. To alleviate this annotation burden…
StyleSentinel: Reliable Artistic Copyright Verification via Stylistic Fingerprints
Lingxiao Chen, Liqin Wang, Wei Lu
The versatility of diffusion models in generating customized images has led to unauthorized usage of personal artwork, which poses a significant threat to the intellectual property…