1 paper
Junbeom Hong, Seonghoon Yu, Hyung Rok Jung +2
Collecting natural-language referring expressions along with region annotations, such as masks or boxes, is a major bottleneck in visual grounding (VG), as annotators must write de…