53 citations · 53 across the 2 of their papers we have counts for
2 papers
cs.CV2023★ 53 cited
Context Disentangling and Prototype Inheriting for Robust Visual Grounding
Wei Tang, Liang Li, Xuejing Liu +3
Visual grounding (VG) aims to locate a specific target in an image based on a given language query. The discriminative information from context is important for distinguishing the…
cs.CL2023
Listener Model for the PhotoBook Referential Game with CLIPScores as Implicit Reference Chain
Shih-Lun Wu, Yi-Hui Chou, Liangze Li
PhotoBook is a collaborative dialogue game where two players receive private, partially-overlapping sets of images and resolve which images they have in common. It presents machine…