Showing cs.CVShow all
2 papers · 1 filter
cs.CV2024
The Solution for the 5th GCAIAC Zero-shot Referring Expression Comprehension Challenge
Longfei Huang, Feng Yu, Zhihao Guan +2
This report presents a solution for the zero-shot referring expression comprehension task. Visual-language multimodal base models (such as CLIP, SAM) have gained significant attent…
cs.CV2024
The Solution for Language-Enhanced Image New Category Discovery
Haonan Xu, Dian Chao, Xiangyu Wu +2
Treating texts as images, combining prompts with textual labels for prompt tuning, and leveraging the alignment properties of CLIP have been successfully applied in zero-shot multi…