1 paper
Seonghoon Yu, Junbeom Hong, Joonseok Lee +1
Visual grounding tasks, such as referring image segmentation (RIS) and referring expression comprehension (REC), aim to localize a target object based on a given textual descriptio…