Showing cs.CVShow all
3 papers · 1 filter
cs.CV2024
SynthDoc: Bilingual Documents Synthesis for Visual Document Understanding
Chuanghao Ding, Xuejing Liu, Wei Tang +5
This paper introduces SynthDoc, a novel synthetic document generation pipeline designed to enhance Visual Document Understanding (VDU) by generating high-quality, diverse datasets…
cs.CV2024
What Makes Good Few-shot Examples for Vision-Language Models?
Zhaojun Guo, Jinghui Lu, Xuejing Liu +3
Despite the notable advancements achieved by leveraging pre-trained vision-language (VL) models through few-shot tuning for downstream tasks, our detailed empirical study highlight…
cs.CV2023
What Large Language Models Bring to Text-rich VQA?
Xuejing Liu, Wei Tang, Xinzhe Ni +4
Text-rich VQA, namely Visual Question Answering based on text recognition in the images, is a cross-modal task that requires both image comprehension and text recognition. In this…