5 papers
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes
Yuji Wang, Haoran Xu, Yong Liu +2
Reference Audio-Visual Segmentation (Ref-AVS) aims to provide a pixel-wise scene understanding in Language-aided Audio-Visual Scenes (LAVS). This task requires the model to continu…
iMOVE: Instance-Motion-Aware Video Understanding
Jiaze Li, Yaya Shi, Zongyang Ma +7
Enhancing the fine-grained instance spatiotemporal motion perception capabilities of Video Large Language Models is crucial for improving their temporal and general video understan…
Federated Learning with Sample-level Client Drift Mitigation
Haoran Xu, Jiaze Li, Wanyi Wu +1
Federated Learning (FL) suffers from severe performance degradation due to the data heterogeneity among clients. Existing works reveal that the fundamental reason is that data hete…
Multilevel Semantic-Aware Model for AI-Generated Video Quality Assessment
Jiaze Li, Haoran Xu, Shiding Zhu +2
The rapid development of diffusion models has greatly advanced AI-generated videos in terms of length and consistency recently, yet assessing AI-generated videos still remains chal…
NTIRE 2024 Challenge on Short-form UGC Video Quality Assessment: Methods and Results
Xin Li, Kun Yuan, Yajing Pei +65
This paper reviews the NTIRE 2024 Challenge on Shortform UGC Video Quality Assessment (S-UGC VQA), where various excellent solutions are submitted and evaluated on the collected da…