Self-Supervised Learning for Text Recognition: A Critical Survey
arXiv:2407.19889 · doi:10.1007/s11263-025-02487-3
Abstract
Text Recognition (TR) refers to the research area that focuses on retrieving textual information from images, a topic that has seen significant advancements in the last decade due to the use of Deep Neural Networks (DNN). However, these solutions often necessitate vast amounts of manually labeled or synthetic data. Addressing this challenge, Self-Supervised Learning (SSL) has gained attention by utilizing large datasets of unlabeled data to train DNN, thereby generating meaningful and robust representations. Although SSL was initially overlooked in TR because of its unique characteristics, recent years have witnessed a surge in the development of SSL methods specifically for this field. This rapid development, however, has led to many methods being explored independently, without taking previous efforts in methodology or comparison into account, thereby hindering progress in the field of research. This paper, therefore, seeks to consolidate the use of SSL in the field of TR, offering a critical and comprehensive overview of the current state of the art. We will review and analyze the existing methods, compare their results, and highlight inconsistencies in the current literature. This thorough analysis aims to provide general insights into the field, propose standardizations, identify new research directions, and foster its proper development.
Published at International Journal of Computer Vision (IJCV)
References in corpus (22)
- Fundamentals of Recurrent Neural Network (RNN) and Long Short-Term Memory (LSTM) Network
- Generative Adversarial Networks: An Overview
- Bootstrap your own latent: A new approach to self-supervised Learning
- Unsupervised Learning of Visual Features by Contrasting Cluster Assignments
- Pre-trained Models for Natural Language Processing: A Survey
- Barlow Twins: Self-Supervised Learning via Redundancy Reduction
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers
- Focusing Attention: Towards Accurate Text Recognition in Natural Images
- Image inpainting: A review
- End-to-end Handwritten Paragraph Text Recognition Using a Vertical Attention Network
- DAN: a Segmentation-free Document Attention Network for Handwritten Document Recognition
- Listen and Fill in the Missing Letters: Non-Autoregressive Transformer for Speech Recognition
- Unsupervised Adaptation for Synthetic-to-Real Handwritten Word Recognition
- ReSSL: Relational Self-Supervised Learning with Weak Augmentation
- Self-Supervised Learning for Data Scarcity in a Fatigue Damage Prognostic Problem
- A Study of the Generalizability of Self-Supervised Representations
- Source Data-absent Unsupervised Domain Adaptation through Hypothesis Transfer and Labeling Transfer
- Efficient Self-supervised Learning with Contextualized Target Representations for Vision, Speech and Language
- Exploring DeshuffleGANs in Self-Supervised Generative Adversarial Networks
- VICRegL: Self-Supervised Learning of Local Visual Features
- Reverse Engineering Self-Supervised Learning
- Image-Text Pre-Training for Logo Recognition