11 papers
VidAudio-Bench: Benchmarking V2A and VT2A Generation across Four Audio Categories
Qian Zhang, Yuqin Cao, Yixuan Gao +1
Video-to-Audio (V2A) generation is essential for immersive multimedia experiences, yet its evaluation remains underexplored. Existing benchmarks typically assess diverse audio type…
Ges-QA: A Multidimensional Quality Assessment Dataset for Audio-to-3D Gesture Generation
Zhilin Gao, Yunhao Li, Sijing Wu +3
The Audio-to-3D-Gesture (A2G) task has enormous potential for various applications in virtual reality and computer graphics, etc. However, current evaluation metrics, such as Fréc…
Surveillance Facial Image Quality Assessment: A Multi-dimensional Dataset and Lightweight Model
Yanwei Jiang, Wei Sun, Yingjie Zhou +8
Surveillance facial images are often captured under unconstrained conditions, resulting in severe quality degradation due to factors such as low resolution, motion blur, occlusion,…
XGC-AVis: Towards Audio-Visual Content Understanding with a Multi-Agent Collaborative System
Yuqin Cao, Xiongkuo Min, Yixuan Gao +4
In this paper, we propose XGC-AVis, a multi-agent framework that enhances the audio-video temporal alignment capabilities of multimodal large models (MLLMs) and improves the effici…
VQualA 2025 Challenge on Engagement Prediction for Short Videos: Methods and Results
Dasong Li, Sizhuo Ma, Hang Hua +40
This paper presents an overview of the VQualA 2025 Challenge on Engagement Prediction for Short Videos, held in conjunction with ICCV 2025. The challenge focuses on understanding a…
Engagement Prediction of Short Videos with Large Multimodal Models
Wei Sun, Linhan Cao, Yuqin Cao +7
The rapid proliferation of user-generated content (UGC) on short-form video platforms has made video engagement prediction increasingly important for optimizing recommendation syst…