1 paper
Xinyu Ma, Ziyang Ding, Zhicong Luo +6
Knowledge-Intensive Visual Grounding (KVG) requires models to localize objects using fine-grained, domain-specific entity names rather than generic referring expressions. Although…