From the 1 of 5 linked papers with an AI index.
5 papers
MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation
Xiaohan Zhang, Yuqing Wen, Junlin Chen +9
The paper introduces MultiRef-Compass, a benchmark for evaluating models that generate synchronized audio‑video content conditioned on multiple references and textual instructions,…
Shape of Thought: Progressive Object Assembly via Visual Chain-of-Thought
Yu Huo, Siyu Zhang, Kun Zeng +7
Multimodal models for text-to-image generation have achieved strong visual fidelity, yet they remain brittle under compositional structural constraints, notably generative numeracy…
VQualA 2025 Challenge on Image Super-Resolution Generated Content Quality Assessment: Methods and Results
Yixiao Li, Xin Li, Chris Wei Zhou +28
This paper presents the ISRGC-Q Challenge, built upon the Image Super-Resolution Generated Content Quality Assessment (ISRGen-QA) dataset, and organized as part of the Visual Quali…
DreamLight: Towards Harmonious and Consistent Image Relighting
Yong Liu, Wenpeng Xiao, Qianqian Wang +5
We introduce a model named DreamLight for universal image relighting in this work, which can seamlessly composite subjects into a new background while maintaining aesthetic uniform…
Tuning-Free Long Video Generation via Global-Local Collaborative Diffusion
Yongjia Ma, Junlin Chen, Donglin Di +6
Creating high-fidelity, coherent long videos is a sought-after aspiration. While recent video diffusion models have shown promising potential, they still grapple with spatiotempora…