33 citations · 215 across the 144 of their papers we have counts for
8 papers · 1 filter
XGC-AVis: Towards Audio-Visual Content Understanding with a Multi-Agent Collaborative System
Yuqin Cao, Xiongkuo Min, Yixuan Gao +4
In this paper, we propose XGC-AVis, a multi-agent framework that enhances the audio-video temporal alignment capabilities of multimodal large models (MLLMs) and improves the effici…
EEmo-Bench: A Benchmark for Multi-modal Large Language Models on Image Evoked Emotion Assessment
Lancheng Gao, Ziheng Jia, Yunhao Zeng +5
The furnishing of multi-modal large language models (MLLMs) has led to the emergence of numerous benchmark studies, particularly those evaluating their perception and understanding…
AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment
Yuqin Cao, Xiongkuo Min, Yixuan Gao +2
Many video-to-audio (VTA) methods have been proposed for dubbing silent AI-generated videos. An efficient quality assessment method for AI-generated audio-visual content (AGAV) is…
Subjective and Objective Quality-of-Experience Evaluation Study for Live Video Streaming
Zehao Zhu, Wei Sun, Jun Jia +7
In recent years, live video streaming has gained widespread popularity across various social media platforms. Quality of experience (QoE), which reflects end-users' satisfaction an…
Towards Effective User Attribution for Latent Diffusion Models via Watermark-Informed Blending
Yongyang Pan, Xiaohong Liu, Siqi Luo +5
Rapid advancements in multimodal large language models have enabled the creation of hyper-realistic images from textual descriptions. However, these advancements also raise signifi…
G-Refine: A General Quality Refiner for Text-to-Image Generation
Chunyi Li, Haoning Wu, Hongkun Hao +7
With the evolution of Text-to-Image (T2I) models, the quality defects of AI-Generated Images (AIGIs) pose a significant barrier to their widespread adoption. In terms of both perce…