3 papers
cs.CV2025
VITAL: Vision-Encoder-centered Pre-training for LMMs in Visual Quality Assessment
Ziheng Jia, Linhan Cao, Jinliang Han +6
Developing a robust visual quality assessment (VQualA) large multi-modal model (LMM) requires achieving versatility, powerfulness, and transferability. However, existing VQualA LMM…
cs.MM2025
XGC-AVis: Towards Audio-Visual Content Understanding with a Multi-Agent Collaborative System
Yuqin Cao, Xiongkuo Min, Yixuan Gao +4
In this paper, we propose XGC-AVis, a multi-agent framework that enhances the audio-video temporal alignment capabilities of multimodal large models (MLLMs) and improves the effici…
cs.CV2025
Scaling-up Perceptual Video Quality Assessment
Ziheng Jia, Zicheng Zhang, Zeyu Zhang +12
The data scaling law has been shown to significantly enhance the performance of large multi-modal models (LMMs) across various downstream tasks. However, in the domain of perceptua…