11 citations · 11 across the 2 of their papers we have counts for
5 papers
Context Disentangling and Prototype Inheriting for Robust Visual Grounding
Wei Tang, Liang Li, Xuejing Liu +3
Visual grounding (VG) aims to locate a specific target in an image based on a given language query. The discriminative information from context is important for distinguishing the…
What Large Language Models Bring to Text-rich VQA?
Xuejing Liu, Wei Tang, Xinzhe Ni +4
Text-rich VQA, namely Visual Question Answering based on text recognition in the images, is a cross-modal task that requires both image comprehension and text recognition. In this…
Parsing-based View-aware Embedding Network for Vehicle Re-Identification
Dechao Meng, Liang Li, Xuejing Liu +6
Vehicle Re-Identification is to find images of the same vehicle from various views in the cross-camera scenario. The main challenges of this task are the large intra-instance dista…
Knowledge-guided Pairwise Reconstruction Network for Weakly Supervised Referring Expression Grounding
Xuejing Liu, Liang Li, Shuhui Wang +3
Weakly supervised referring expression grounding (REG) aims at localizing the referential entity in an image according to linguistic query, where the mapping between the image regi…
Adaptive Reconstruction Network for Weakly Supervised Referring Expression Grounding
Xuejing Liu, Liang Li, Shuhui Wang +3
Weakly supervised referring expression grounding aims at localizing the referential object in an image according to the linguistic query, where the mapping between the referential…