3 citations · 4 across the 2 of their papers we have counts for
2 papers
cs.CL2023★ 1 cited
A Multi-Modal Multilingual Benchmark for Document Image Classification
Yoshinari Fujinuma, Siddharth Varia, Nishant Sankaran +3
Document image classification is different from plain-text document classification and consists of classifying a document by understanding the content and structure of documents su…
cs.CV2023★ 3 cited
DocFormerv2: Local Features for Document Understanding
Srikar Appalaraju, Peng Tang, Qi Dong +3
We propose DocFormerv2, a multi-modal transformer for Visual Document Understanding (VDU). The VDU domain entails understanding documents (beyond mere OCR predictions) e.g., extrac…