15 papers
SpongeBob: Sync-Aware Harmonious Audio-Visual Generative Editing
Sen Liang, Cong Wang, Fengbin Guan +6
Visual and acoustic events in the physical world are inherently coupled, yet existing video editing methods typically adopt decoupled pipelines, lacking bidirectional modality inte…
IQA-Spider: Unifying Multi-Granularity Image Quality Assessment with Reasoning, Grounding and Referring
Xinge Peng, Yiting Lu, Xin Li +1
We present IQA-Spider, the first image quality assessment (IQA) framework that unifies reasoning, grounding, and referring into a single LMM-based framework for multi-granularity q…
LoViF 2026 The First Challenge on Holistic Quality Assessment for 4D World Model (PhyScore)
Wei Luo, Yiting Lu, Xin Li +32
This paper reports on the LoViF 2026 PhyScore challenge, a competition on holistic quality assessment of world-model-generated videos across both 2D and 4D generation settings. The…
DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning
Yuancheng Wei, Haojie Zhang, Linli Yao +7
Image Difference Captioning (IDC) generates natural language descriptions that precisely identify differences between two images, serving as a key benchmark for fine-grained change…
4DWorldBench: A Comprehensive Evaluation Framework for 3D/4D World Generation Models
Yiting Lu, Wei Luo, Peiyan Tu +8
World Generation Models are emerging as a cornerstone of next-generation multimodal intelligence systems. Unlike traditional 2D visual generation, World Models aim to construct rea…
OmniQuality-R: Advancing Reward Models Through All-Encompassing Quality Assessment
Yiting Lu, Fengbin Guan, Yixin Gao +8
Current visual evaluation approaches are typically constrained to a single task. To address this, we propose OmniQuality-R, a unified reward modeling framework that transforms mult…