3 papers
cs.MM2025
When Harmful Content Gets Camouflaged: Unveiling Perception Failure of LVLMs with CamHarmTI
Yanhui Li, Qi Zhou, Zhihong Xu +3
Large vision-language models (LVLMs) are increasingly used for tasks where detecting multimodal harmful content is crucial, such as online content moderation. However, real-world h…
cs.MM2025
Sonic: Shifting Focus to Global Audio Perception in Portrait Animation
Xiaozhong Ji, Xiaobin Hu, Zhihong Xu +9
The study of talking face generation mainly explores the intricacies of synchronizing facial movements and crafting visually appealing, temporally-coherent animations. However, due…
cs.CV2025
TruePose: Human-Parsing-guided Attention Diffusion for Full-ID Preserving Pose Transfer
Zhihong Xu, Dongxia Wang, Peng Du +2
Pose-Guided Person Image Synthesis (PGPIS) generates images that maintain a subject's identity from a source image while adopting a specified target pose (e.g., skeleton). While di…