5 papers · 1 filter
ITIScore: An Image-to-Text-to-Image Rating Framework for the Image Captioning Ability of MLLMs
Zitong Xu, Huiyu Duan, Shengyao Qin +6
Recent advances in multimodal large language models (MLLMs) have greatly improved image understanding and captioning capabilities. However, existing image captioning benchmarks typ…
Robust Mesh Saliency Ground Truth Acquisition in VR via View Cone Sampling and Manifold Diffusion
Guoquan Zheng, Jie Hao, Huiyu Duan +7
As the complexity of 3D digital content grows exponentially, understanding human visual attention is critical for optimizing rendering and processing resources. Therefore, reliable…
Quality Assessment and Distortion-aware Saliency Prediction for AI-Generated Omnidirectional Images
Liu Yang, Huiyu Duan, Jiarui Wang +5
With the rapid advancement of Artificial Intelligence Generated Content (AIGC) techniques, AI generated images (AIGIs) have attracted widespread attention, among which AI generated…
Efficient Listener: Dyadic Facial Motion Synthesis via Action Diffusion
Zesheng Wang, Alexandre Bruckert, Patrick Le Callet +1
Generating realistic listener facial motions in dyadic conversations remains challenging due to the high-dimensional action space and temporal dependency requirements. Existing app…
ESVQA: Perceptual Quality Assessment of Egocentric Spatial Videos
Xilei Zhu, Huiyu Duan, Liu Yang +4
With the rapid development of eXtended Reality (XR), egocentric spatial shooting and display technologies have further enhanced immersion and engagement for users, delivering more…