1 paper
Longfei Huang, Feng Yu, Zhihao Guan +2
This report presents a solution for the zero-shot referring expression comprehension task. Visual-language multimodal base models (such as CLIP, SAM) have gained significant attent…