2 papers
cs.CV2024
The Solution for the 5th GCAIAC Zero-shot Referring Expression Comprehension Challenge
Longfei Huang, Feng Yu, Zhihao Guan +2
This report presents a solution for the zero-shot referring expression comprehension task. Visual-language multimodal base models (such as CLIP, SAM) have gained significant attent…
cs.CV2024
The Solution for the ICCV 2023 Perception Test Challenge 2023 -- Task 6 -- Grounded videoQA
Hailiang Zhang, Dian Chao, Zhihao Guan +1
In this paper, we introduce a grounded video question-answering solution. Our research reveals that the fixed official baseline method for video question answering involves two mai…