1 paper
Viet-Quoc Pham, Nao Mishima
Weakly supervised visual grounding aims to predict the region in an image that corresponds to a specific linguistic query, where the mapping between the target object and query is…