2 citations · 2 across the 10 of their papers we have counts for
8 papers · 1 filter
Towards Explainable Fake Image Detection with Multi-Modal Large Language Models
Yikun Ji, Yan Hong, Jiahui Zhan +6
Progress in image generation raises significant public security concerns. We argue that fake image detection should not operate as a "black box". Instead, an ideal approach must en…
DS-VTON: An Enhanced Dual-Scale Coarse-to-Fine Framework for Virtual Try-On
Xianbing Sun, Yan Hong, Jiahui Zhan +5
Despite recent progress, most existing virtual try-on methods still struggle to simultaneously address two core challenges: accurately aligning the garment image with the target hu…
InterAnimate: Taming Region-aware Diffusion Model for Realistic Human Interaction Animation
Yukang Lin, Yan Hong, Zunnan Xu +10
Recent video generation research has focused heavily on isolated actions, leaving interactive motions-such as hand-face interactions-largely unexamined. These interactions are esse…
EMIT: Enhancing MLLMs for Industrial Anomaly Detection via Difficulty-Aware GRPO
Wei Guan, Jun Lan, Jian Cao +3
Industrial anomaly detection (IAD) plays a crucial role in maintaining the safety and reliability of manufacturing systems. While multimodal large language models (MLLMs) show stro…
Interpretable and Reliable Detection of AI-Generated Images via Grounded Reasoning in MLLMs
Yikun Ji, Hong Yan, Jun Lan +5
The rapid advancement of image generation technologies intensifies the demand for interpretable and robust detection methods. Although existing approaches often attain high accurac…
Stochastic Layer-Wise Shuffle for Improving Vision Mamba Training
Zizheng Huang, Haoxing Chen, Jiaqi Li +4
Recent Vision Mamba (Vim) models exhibit nearly linear complexity in sequence length, making them highly attractive for processing visual data. However, the training methodologies…