1 paper
Yichi Zhang, Gongwei Chen, Jun Zhu +2
Visual grounding requires large and diverse region-text pairs. However, manual annotation is costly and fixed vocabularies restrict scalability and generalization. Existing pseudo-…