11 papers
GLaVE-Cap: Global-Local Aligned Video Captioning with Vision Expert Integration
Wan Xu, Feng Zhu, Yihan Zeng +4
Video detailed captioning aims to generate comprehensive video descriptions to facilitate video understanding. Recently, most efforts in the video detailed captioning community hav…
MDIQA: Unified Image Quality Assessment for Multi-dimensional Evaluation and Restoration
Shunyu Yao, Ming Liu, Zhilu Zhang +4
Recent advancements in image quality assessment (IQA), driven by sophisticated deep neural network designs, have significantly improved the ability to approach human perceptions. H…
RefSTAR: Blind Facial Image Restoration with Reference Selection, Transfer, and Reconstruction
Zhicun Yin, Junjie Chen, Ming Liu +6
Blind facial image restoration is highly challenging due to unknown complex degradations and the sensitivity of humans to faces. Although existing methods introduce auxiliary infor…
Integrating Visual Interpretation and Linguistic Reasoning for Math Problem Solving
Zixian Guo, Ming Liu, Qilong Wang +4
Current large vision-language models (LVLMs) typically employ a connector module to link visual features with text embeddings of large language models (LLMs) and use end-to-end tra…
NTIRE 2025 Challenge on Real-World Face Restoration: Methods and Results
Zheng Chen, Jingkai Wang, Kai Liu +51
This paper provides a review of the NTIRE 2025 challenge on real-world face restoration, highlighting the proposed solutions and the resulting outcomes. The challenge focuses on ge…
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection
Weijun Zhuang, Qizhang Li, Xin Li +5
Temporal Action Detection and Moment Retrieval constitute two pivotal tasks in video understanding, focusing on precisely localizing temporal segments corresponding to specific act…