4 papers
The Solution for the CVPR2024 NICE Image Captioning Challenge
Longfei Huang, Shupeng Zhong, Xiangyu Wu +1
This report introduces a solution to the Topic 1 Zero-shot Image Captioning of 2024 NICE : New frontiers for zero-shot Image Captioning Evaluation. In contrast to NICE 2023 dataset…
The Solution for the ICCV 2023 1st Scientific Figure Captioning Challenge
Dian Chao, Xin Song, Shupeng Zhong +4
In this paper, we propose a solution for improving the quality of captions generated for figures in papers. We adopt the approach of summarizing the textual content in the paper to…
Solution for SMART-101 Challenge of ICCV Multi-modal Algorithmic Reasoning Task 2023
Xiangyu Wu, Yang Yang, Shengdong Xu +3
In this paper, we present our solution to a Multi-modal Algorithmic Reasoning Task: SMART-101 Challenge. Different from the traditional visual question-answering datasets, this cha…
Generation-Guided Multi-Level Unified Network for Video Grounding
Xing Cheng, Xiangyu Wu, Dong Shen +2
Video grounding aims to locate the timestamps best matching the query description within an untrimmed video. Prevalent methods can be divided into moment-level and clip-level frame…