2 papers
cs.CV2026
A Proposal-Free Query-Guided Network for Grounded Multimodal Named Entity Recognition
Hongbing Li, Jiamin Liu, Shuo Zhang +1
Grounded Multimodal Named Entity Recognition (GMNER) identifies named entities, including their spans and types, in natural language text and grounds them to the corresponding regi…
cs.CL2024
Enhancing Question Answering on Charts Through Effective Pre-training Tasks
Ashim Gupta, Vivek Gupta, Shuo Zhang +3
To completely understand a document, the use of textual information is not enough. Understanding visual cues, such as layouts and charts, is also required. While the current state-…