collaborators

5 papers

cs.CV2025

GRADEO: Towards Human-Like Evaluation for Text-to-Video Generation via Multi-Step Reasoning

Zhun Mou, Bin Xia, Zhengchao Huang +2

Recent great advances in video generation models have demonstrated their potential to produce high-quality videos, bringing challenges to effective evaluation. Unlike human evaluat…

cs.CV2025

CoordSpeaker: Exploiting Gesture Captioning for Coordinated Caption-Empowered Co-Speech Gesture Generation

Fengyi Fang, Sicheng Yang, Wenming Yang

Co-speech gesture generation has significantly advanced human-computer interaction, yet speaker movements remain constrained due to the omission of text-driven non-spontaneous gest…

cs.CV2025

Alignment is All You Need: A Training-free Augmentation Strategy for Pose-guided Video Generation

Xiaoyu Jin, Zunnan Xu, Mingwen Ou +1

Character animation is a transformative field in computer graphics and vision, enabling dynamic and realistic video animations from static images. Despite advancements, maintaining…

cs.CV2025

DiffStereo: High-Frequency Aware Diffusion Model for Stereo Image Restoration

Huiyun Cao, Yuan Shi, Bin Xia +2

Diffusion models (DMs) have achieved promising performance in image restoration but haven't been explored for stereo images. The application of DM in stereo image restoration is co…

cs.CV2024

FFAA: Multimodal Large Language Model based Explainable Open-World Face Forgery Analysis Assistant

Zhengchao Huang, Bin Xia, Zicheng Lin +3

The rapid advancement of deepfake technologies has sparked widespread public concern, particularly as face forgery poses a serious threat to public information security. However, t…