3 papers
cs.CV2024
SynthDoc: Bilingual Documents Synthesis for Visual Document Understanding
Chuanghao Ding, Xuejing Liu, Wei Tang +5
This paper introduces SynthDoc, a novel synthetic document generation pipeline designed to enhance Visual Document Understanding (VDU) by generating high-quality, diverse datasets…
cs.CV2023
What Large Language Models Bring to Text-rich VQA?
Xuejing Liu, Wei Tang, Xinzhe Ni +4
Text-rich VQA, namely Visual Question Answering based on text recognition in the images, is a cross-modal task that requires both image comprehension and text recognition. In this…
cs.CV2022
Entity-enhanced Adaptive Reconstruction Network for Weakly Supervised Referring Expression Grounding
Xuejing Liu, Liang Li, Shuhui Wang +4
Weakly supervised Referring Expression Grounding (REG) aims to ground a particular target in an image described by a language expression while lacking the correspondence between ta…