1 paper
Minghang Zheng, Jiahua Zhang, Qingchao Chen +2
Visual grounding aims to localize the object referred to in an image based on a natural language query. Although progress has been made recently, accurately localizing target objec…