activity
20242026
most citedTowards Explainable Fake Image Detection with Multi-Modal Large Language Models

2 citations · 2 across the 10 of their papers we have counts for

collaborators
Showing 2025Show all

8 papers · 1 filter

cs.CV20252 cited

Towards Explainable Fake Image Detection with Multi-Modal Large Language Models

Yikun Ji, Yan Hong, Jiahui Zhan +6

Progress in image generation raises significant public security concerns. We argue that fake image detection should not operate as a "black box". Instead, an ideal approach must en…

cs.CV2025

DS-VTON: An Enhanced Dual-Scale Coarse-to-Fine Framework for Virtual Try-On

Xianbing Sun, Yan Hong, Jiahui Zhan +5

Despite recent progress, most existing virtual try-on methods still struggle to simultaneously address two core challenges: accurately aligning the garment image with the target hu…

cs.CV2025

InterAnimate: Taming Region-aware Diffusion Model for Realistic Human Interaction Animation

Yukang Lin, Yan Hong, Zunnan Xu +10

Recent video generation research has focused heavily on isolated actions, leaving interactive motions-such as hand-face interactions-largely unexamined. These interactions are esse…

cs.CV2025

EMIT: Enhancing MLLMs for Industrial Anomaly Detection via Difficulty-Aware GRPO

Wei Guan, Jun Lan, Jian Cao +3

Industrial anomaly detection (IAD) plays a crucial role in maintaining the safety and reliability of manufacturing systems. While multimodal large language models (MLLMs) show stro…

cs.CV2025

Interpretable and Reliable Detection of AI-Generated Images via Grounded Reasoning in MLLMs

Yikun Ji, Hong Yan, Jun Lan +5

The rapid advancement of image generation technologies intensifies the demand for interpretable and robust detection methods. Although existing approaches often attain high accurac…

cs.CV2025

Stochastic Layer-Wise Shuffle for Improving Vision Mamba Training

Zizheng Huang, Haoxing Chen, Jiaqi Li +4

Recent Vision Mamba (Vim) models exhibit nearly linear complexity in sequence length, making them highly attractive for processing visual data. However, the training methodologies…