3 citations · 6 across the 3 of their papers we have counts for
1 paper · 1 filter
Anwen Hu, Haiyang Xu, Jiabo Ye +8
Structure information is critical for understanding the semantics of text-rich images, such as documents, tables, and charts. Existing Multimodal Large Language Models (MLLMs) for…