11 papers · 1 filter
Traffic Sign Recognition in Autonomous Driving: Dataset, Benchmark, and Field Experiment
Guoyang Zhao, Weiqing Qi, Kai Zhang +7
Traffic Sign Recognition (TSR) is a core perception capability for autonomous driving, where robustness to cross-region variation, long-tailed categories, and semantic ambiguity is…
Cluster-Wise Spatio-Temporal Masking for Efficient Video-Language Pretraining
Weijun Zhuang, Yuqing Huang, Weikang Meng +5
Large-scale video-language pretraining enables strong generalization across multimodal tasks but often incurs prohibitive computational costs. Although recent advances in masked vi…
GLaVE-Cap: Global-Local Aligned Video Captioning with Vision Expert Integration
Wan Xu, Feng Zhu, Yihan Zeng +4
Video detailed captioning aims to generate comprehensive video descriptions to facilitate video understanding. Recently, most efforts in the video detailed captioning community hav…
MDIQA: Unified Image Quality Assessment for Multi-dimensional Evaluation and Restoration
Shunyu Yao, Ming Liu, Zhilu Zhang +4
Recent advancements in image quality assessment (IQA), driven by sophisticated deep neural network designs, have significantly improved the ability to approach human perceptions. H…
RefSTAR: Blind Facial Image Restoration with Reference Selection, Transfer, and Reconstruction
Zhicun Yin, Junjie Chen, Ming Liu +6
Blind facial image restoration is highly challenging due to unknown complex degradations and the sensitivity of humans to faces. Although existing methods introduce auxiliary infor…
NTIRE 2025 Challenge on Real-World Face Restoration: Methods and Results
Zheng Chen, Jingkai Wang, Kai Liu +51
This paper provides a review of the NTIRE 2025 challenge on real-world face restoration, highlighting the proposed solutions and the resulting outcomes. The challenge focuses on ge…