CLIP4STR: A Simple Baseline for Scene Text Recognition with Pre-trained Vision-Language Model
arXiv:2305.14014 · doi:10.1109/TIP.2024.3512354
Abstract
Pre-trained vision-language models~(VLMs) are the de-facto foundation models for various downstream tasks. However, scene text recognition methods still prefer backbones pre-trained on a single modality, namely, the visual modality, despite the potential of VLMs to serve as powerful scene text readers. For example, CLIP can robustly identify regular (horizontal) and irregular (rotated, curved, blurred, or occluded) text in images. With such merits, we transform CLIP into a scene text reader and introduce CLIP4STR, a simple yet effective STR method built upon image and text encoders of CLIP. It has two encoder-decoder branches: a visual branch and a cross-modal branch. The visual branch provides an initial prediction based on the visual feature, and the cross-modal branch refines this prediction by addressing the discrepancy between the visual feature and text semantics. To fully leverage the capabilities of both branches, we design a dual predict-and-refine decoding scheme for inference. We scale CLIP4STR in terms of the model size, pre-training data, and training data, achieving state-of-the-art performance on 13 STR benchmarks. Additionally, a comprehensive empirical study is provided to enhance the understanding of the adaptation of CLIP to STR. Our method establishes a simple yet strong baseline for future STR research with VLMs.
Accepted by T-IP. A PyTorch re-implementation is at https://github.com/VamosC/CLIP4STR (Credit on GitHub@VamosC)
References in corpus (17)
- Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization
- Learning to Prompt for Vision-Language Models
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- Flamingo: a Visual Language Model for Few-Shot Learning
- Synthetic Data and Artificial Neural Networks for Natural Scene Text Recognition
- Focusing Attention: Towards Accurate Text Recognition in Natural Images
- CoCa: Contrastive Captioners are Image-Text Foundation Models
- Florence: A New Foundation Model for Computer Vision
- Rosetta: Large scale system for text detection and recognition in images
- Towards artificial general intelligence via a multimodal foundation model
- FILIP: Fine-grained Interactive Language-Image Pre-Training
- CenterCLIP: Token Clustering for Efficient Text-Video Retrieval
- Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm
- ERNIE-ViLG: Unified Generative Pre-training for Bidirectional Vision-Language Generation
- Data Filtering Networks
- Masked and Permuted Implicit Context Learning for Scene Text Recognition
- Scene Text Recognition Models Explainability Using Local Features
Cited by in corpus (5)
- An Explainable Deep Neural Network with Frequency-Aware Channel and Spatial Refinement for Flood Prediction in Sustainable Cities
- SELECT: Detecting Label Errors in Real-world Scene Text Data
- Bharat Scene Text: A Novel Comprehensive Dataset and Benchmark for Indian Language Scene Text Understanding
- Stratified Domain Adaptation: A Progressive Self-Training Approach for Scene Text Recognition
- HAAP: Vision-context Hierarchical Attention Autoregressive with Adaptive Permutation for Scene Text Recognition