3 papers
cs.CV2023
Do-GOOD: Towards Distribution Shift Evaluation for Pre-Trained Visual Document Understanding Models
Jiabang He, Yi Hu, Lei Wang +4
Numerous pre-training techniques for visual document understanding (VDU) have recently shown substantial improvements in performance across a wide range of document tasks. However,…
cs.CL2023
T-SciQ: Teaching Multimodal Chain-of-Thought Reasoning via Mixed Large Language Model Signals for Science Question Answering
Lei Wang, Yi Hu, Jiabang He +4
Large Language Models (LLMs) have recently demonstrated exceptional performance in various Natural Language Processing (NLP) tasks. They have also shown the ability to perform chai…
cs.CV2022
Alignment-Enriched Tuning for Patch-Level Pre-trained Document Image Models
Lei Wang, Jiabang He, Xing Xu +2
Alignment between image and text has shown promising improvements on patch-level pre-trained document image models. However, investigating more effective or finer-grained alignment…