From the 1 of 28 linked papers with an AI index.
1 citations · 2 across the 10 of their papers we have counts for
15 papers · 1 filter
UniHetero: Could Generation Enhance Understanding for Vision-Language-Model at Large Data Scale?
Fengjiao Chen, Minhao Jing, Weitao Lu +3
Vision-language large models are moving toward the unification of visual understanding and visual generation tasks. However, whether generation can enhance understanding is still u…
VoiceCloak: A Multi-Dimensional Defense Framework against Unauthorized Diffusion-based Voice Cloning
Qianyue Hu, Junyan Wu, Wei Lu +1
Diffusion Models (DMs) have achieved remarkable success in realistic voice cloning (VC), while they also increase the risk of malicious misuse. Existing proactive defenses designed…
Fine-grained Multiple Supervisory Network for Multi-modal Manipulation Detecting and Grounding
Xinquan Yu, Wei Lu, Xiangyang Luo
The task of Detecting and Grounding Multi-Modal Media Manipulation (DGM) is a branch of misinformation detection. Unlike traditional binary classification, it includes complex…
Weakly Supervised Multimodal Temporal Forgery Localization via Multitask Learning
Wenbo Xu, Wei Lu, Xiangyang Luo
The spread of Deepfake videos has caused a trust crisis and impaired social stability. Although numerous approaches have been proposed to address the challenges of Deepfake detecti…
A Multimodal Deviation Perceiving Framework for Weakly-Supervised Temporal Forgery Localization
Wenbo Xu, Junyan Wu, Wei Lu +2
Current researches on Deepfake forensics often treat detection as a classification task or temporal forgery localization problem, which are usually restrictive, time-consuming, and…
Weakly-Supervised Image Forgery Localization via Vision-Language Collaborative Reasoning Framework
Ziqi Sheng, Junyan Wu, Wei Lu +1
Image forgery localization aims to precisely identify tampered regions within images, but it commonly depends on costly pixel-level annotations. To alleviate this annotation burden…