6 papers
Audio-Visual Cross-Modal Compression for Generative Face Video Coding
Youmin Xu, Mengxi Guo, Shijie Zhao +4
Generative face video coding (GFVC) is vital for modern applications like video conferencing, yet existing methods primarily focus on video motion while neglecting the significant…
UARE: A Unified Vision-Language Model for Image Quality Assessment, Restoration, and Enhancement
Weiqi Li, Xuanyu Zhang, Bin Chen +7
Image quality assessment (IQA) and image restoration are fundamental problems in low-level vision. Although IQA and restoration are closely connected conceptually, most existing wo…
VQ-Insight: Teaching VLMs for AI-Generated Video Quality Understanding via Progressive Visual Reinforcement Learning
Xuanyu Zhang, Weiqi Li, Shijie Zhao +3
Recent advances in AI-generated content (AIGC) have led to the emergence of powerful text-to-video generation models. Despite these successes, evaluating the quality of AIGC-genera…
MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection
Shuyu Wang, Weiqi Li, Qian Wang +2
Recent advances in AI-generated content (AIGC) have significantly accelerated image editing techniques, driving increasing demand for diverse and fine-grained edits. Despite these…
Q-Insight: Understanding Image Quality via Visual Reinforcement Learning
Weiqi Li, Xuanyu Zhang, Shijie Zhao +4
Image quality assessment (IQA) focuses on the perceptual visual quality of images, playing a crucial role in downstream tasks such as image reconstruction, compression, and generat…
OmniDrag: Enabling Motion Control for Omnidirectional Image-to-Video Generation
Weiqi Li, Shijie Zhao, Chong Mou +6
As virtual reality gains popularity, the demand for controllable creation of immersive and dynamic omnidirectional videos (ODVs) is increasing. While previous text-to-ODV generatio…