15 papers
FMReward: Aligning and Evaluating Audio-Driven 3D Facial Animation with Human Preferences
Sijing Wu, Yunhao Li, Zhilin Gao +4
Audio-driven 3D facial animation is essential for advancing immersion and interactivity in virtual experiences. Although recent advances have shown promising capabilities, the trai…
Global Attention-Fused Image Cropping with Attention-Guided and Global-Aligned Crop Evaluator
Haotian Yang, Zhile Yang, Kin-Man Lam +2
Image cropping aims to improve image aesthetics by preserving important content within an appropriately composed region. However, most existing methods focus primarily on salient r…
FreeShadow: Training-Free Shadow Removal via Illumination Transfer and Selective Content Preservation in Diffusion Models
Yinan Wang, Yan Huang, Yong Xu +1
FreeShadow removes shadows from images without any training by leveraging pretrained diffusion models, using illumination transfer attention to bring lighting cues from non‑shadow…
Multi-Dimensional Quality Assessment for AI-Generated Human-Centric Videos: Dataset and Model
Sijing Wu, Yunhao Li, Huiyu Duan +4
AI-generated human-centric videos play a crucial role in a wide range of modern applications. However, they often suffer from quality issues and semantic mismatches, underscoring t…
ICME 2026 Grand Challenge on Cross-Scenario Defect Detection and Fine-Grained Severity Grading for High-Precision Manufacturing
Wei Sun, Weixia Zhang, Linhan Cao +30
This paper presents the IEEE International Conference on Multimedia and Expo (ICME) 2026 Grand Challenge on Cross-Scenario Defect Detection and Fine-Grained Severity Grading for Hi…
ITIScore: An Image-to-Text-to-Image Rating Framework for the Image Captioning Ability of MLLMs
Zitong Xu, Huiyu Duan, Shengyao Qin +6
Recent advances in multimodal large language models (MLLMs) have greatly improved image understanding and captioning capabilities. However, existing image captioning benchmarks typ…