activity
20232026
most citedGramformer: Learning Crowd Counting via Graph-Modulated Transformer

2 citations · 5 across the 11 of their papers we have counts for

collaborators
Showing cs.CVShow all

13 papers · 1 filter

cs.CV2026

PATE-Forensics: Perception-as-Tool for Explainable Deepfake Forensics with General-Purpose MLLMs

Yaqi Li, Jielun Peng, Yabin Wang +2

Existing explainable deepfake forensic methods typically rely on task-adapted MLLM to jointly address detection, localization, and explanation. Inspired by agent-style tool use, we…

cs.CV2026

A Benchmark for Semi-supervised Multi-modal Crowd Counting

Haoliang Meng, Xiaopeng Hong, Yabin Wang +1

This paper constructs the first benchmark on semi-supervised multi-modal crowd counting. To lay the foundation for this unexplored task, we first formulate the semi-supervised mult…

cs.CV2026

Leave No Stone Unturned: Uncovering Holistic Audio-Visual Intrinsic Coherence for Deepfake Detection

Jielun Peng, Yabin Wang, Yaqi Li +2

The rapid progress of generative AI has enabled hyper-realistic audio-visual deepfakes, intensifying threats to personal security and social trust. Most existing deepfake detectors…

cs.CV20241 cited

ComprehendEdit: A Comprehensive Dataset and Evaluation Framework for Multimodal Knowledge Editing

Yaohui Ma, Xiaopeng Hong, Shizhou Zhang +4

Large multimodal language models (MLLMs) have revolutionized natural language processing and visual understanding, but often contain outdated or inaccurate information. Current mul…

cs.CV2024

Penny-Wise and Pound-Foolish in Deepfake Detection

Yabin Wang, Zhiwu Huang, Su Zhou +2

The diffusion of deepfake technologies has sparked serious concerns about its potential misuse across various domains, prompting the urgent need for robust detection methods. Despi…

cs.CV2024

Multi-modal Crowd Counting via Modal Emulation

Chenhao Wang, Xiaopeng Hong, Zhiheng Ma +3

Multi-modal crowd counting is a crucial task that uses multi-modal cues to estimate the number of people in crowded scenes. To overcome the gap between different modalities, we pro…