2 citations · 5 across the 11 of their papers we have counts for
13 papers · 1 filter
PATE-Forensics: Perception-as-Tool for Explainable Deepfake Forensics with General-Purpose MLLMs
Yaqi Li, Jielun Peng, Yabin Wang +2
Existing explainable deepfake forensic methods typically rely on task-adapted MLLM to jointly address detection, localization, and explanation. Inspired by agent-style tool use, we…
A Benchmark for Semi-supervised Multi-modal Crowd Counting
Haoliang Meng, Xiaopeng Hong, Yabin Wang +1
This paper constructs the first benchmark on semi-supervised multi-modal crowd counting. To lay the foundation for this unexplored task, we first formulate the semi-supervised mult…
Leave No Stone Unturned: Uncovering Holistic Audio-Visual Intrinsic Coherence for Deepfake Detection
Jielun Peng, Yabin Wang, Yaqi Li +2
The rapid progress of generative AI has enabled hyper-realistic audio-visual deepfakes, intensifying threats to personal security and social trust. Most existing deepfake detectors…
ComprehendEdit: A Comprehensive Dataset and Evaluation Framework for Multimodal Knowledge Editing
Yaohui Ma, Xiaopeng Hong, Shizhou Zhang +4
Large multimodal language models (MLLMs) have revolutionized natural language processing and visual understanding, but often contain outdated or inaccurate information. Current mul…
Penny-Wise and Pound-Foolish in Deepfake Detection
Yabin Wang, Zhiwu Huang, Su Zhou +2
The diffusion of deepfake technologies has sparked serious concerns about its potential misuse across various domains, prompting the urgent need for robust detection methods. Despi…
Multi-modal Crowd Counting via Modal Emulation
Chenhao Wang, Xiaopeng Hong, Zhiheng Ma +3
Multi-modal crowd counting is a crucial task that uses multi-modal cues to estimate the number of people in crowded scenes. To overcome the gap between different modalities, we pro…