OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
arXiv:2305.07895 · doi:10.1007/s11432-024-4235-6
Abstract
Large models have recently played a dominant role in natural language processing and multimodal vision-language learning. However, their effectiveness in text-related visual tasks remains relatively unexplored. In this paper, we conducted a comprehensive evaluation of Large Multimodal Models, such as GPT4V and Gemini, in various text-related visual tasks including Text Recognition, Scene Text-Centric Visual Question Answering (VQA), Document-Oriented VQA, Key Information Extraction (KIE), and Handwritten Mathematical Expression Recognition (HMER). To facilitate the assessment of Optical Character Recognition (OCR) capabilities in Large Multimodal Models, we propose OCRBench, a comprehensive evaluation benchmark. OCRBench contains 29 datasets, making it the most comprehensive OCR evaluation benchmark available. Furthermore, our study reveals both the strengths and weaknesses of these models, particularly in handling multilingual text, handwritten text, non-semantic text, and mathematical expression recognition. Most importantly, the baseline results presented in this study could provide a foundational framework for the conception and assessment of innovative strategies targeted at enhancing zero-shot multimodal techniques. The evaluation pipeline and benchmark are available at https://github.com/Yuliang-Liu/MultimodalOCR.
References in corpus (40)
- Learning Transferable Visual Models From Natural Language Supervision
- LLaMA: Open and Efficient Foundation Language Models
- Flamingo: a Visual Language Model for Few-Shot Learning
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
- ICDAR2019 Competition on Scanned Receipt OCR and Information Extraction
- Towards Expert-Level Medical Question Answering with Large Language Models
- LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day
- GIT: A Generative Image-to-text Transformer for Vision and Language
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
- MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
- ELEVATER: A Benchmark and Toolkit for Evaluating Language-Augmented Visual Models
- Evaluating ChatGPT's Information Extraction Capabilities: An Assessment of Performance, Explainability, Calibration, and Faithfulness
- OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents
- DeepSeek-VL: Towards Real-World Vision-Language Understanding
- LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
- Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
- What matters when building vision-language models?
- How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
- On the General Value of Evidence, and Bilingual Scene-Text Visual Question Answering
- ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
- BLIVA: A Simple Multimodal LLM for Better Handling of Text-Rich Visual Questions
- TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
- InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
- Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models
- From Two to One: A New Scene Text Recognizer with Visual Language Modeling Network
- UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding
- DocPedia: Unleashing the Power of Large Multimodal Model in the Frequency Domain for Versatile Document Understanding
- ICL-D3IE: In-Context Learning with Diverse Demonstrations Updating for Document Information Extraction
- InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
- DePlot: One-shot visual language reasoning by plot-to-table translation
- ERNIE-Layout: Layout Knowledge Enhanced Pre-training for Visually-rich Document Understanding
- Ovis: Structural Embedding Alignment for Multimodal Large Language Model
- When Counting Meets HMER: Counting-Aware Network for Handwritten Mathematical Expression Recognition
- Winner Team Mia at TextVQA Challenge 2021: Vision-and-Language Representation Learning with Pre-trained Sequence-to-Sequence Model
- Toward Understanding WordArt: Corner-Guided Transformer for Scene Text Recognition
- Visual Information Extraction in the Wild: Practical Dataset and End-to-end Solution
Cited by in corpus (5)
- Efficient Multimodal Large Language Models: A Survey
- Task-Oriented Semantic Communication in Large Multimodal Models-based Vehicle Networks
- ICDAR 2025 Competition on End-to-End Document Image Machine Translation Towards Complex Layouts
- Robustness of Structured Data Extraction from Perspectively Distorted Documents
- A Survey of OCR Evaluation Methods and Metrics and the Invisibility of Historical Documents