1 paper
Jielong Tang, Xujie Yuan, Jiayang Liu +6
Grounded Multimodal Named Entity Recognition (GMNER) aims to extract named entities and localize their visual regions within image-text pairs, serving as a pivotal capability for v…