Named Entity Recognition and Classification on Historical Documents: A Survey
arXiv:2109.11406 · doi:10.1145/3604931
Abstract
After decades of massive digitisation, an unprecedented amount of historical documents is available in digital format, along with their machine-readable texts. While this represents a major step forward with respect to preservation and accessibility, it also opens up new opportunities in terms of content mining and the next fundamental challenge is to develop appropriate technologies to efficiently search, retrieve and explore information from this 'big data of the past'. Among semantic indexing opportunities, the recognition and classification of named entities are in great demand among humanities scholars. Yet, named entity recognition (NER) systems are heavily challenged with diverse, historical and noisy inputs. In this survey, we present the array of challenges posed by historical documents to NER, inventory existing resources, describe the main approaches deployed so far, and identify key priorities for future developments.
39 pages
References in corpus (7)
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Natural Language Processing (almost) from Scratch
- A Survey on Recent Advances in Named Entity Recognition from Deep Learning models
- Few-shot classification in Named Entity Recognition Task
- Latin BERT: A Contextual Language Model for Classical Philology
- Digital interfaces of historical newspapers: opportunities, restrictions and recommendations
- hmBERT: Historical Multilingual Language Models for Named Entity Recognition
Cited by in corpus (4)
- ChroniclingAmericaQA: A Large-scale Question Answering Dataset based on Historical American Newspaper Pages
- Musical Heritage Historical Entity Linking
- Is text normalization relevant for classifying medieval charters?
- ICDAR 2026 HIPE-OCRepair Competition on LLM-Assisted OCR Post-Correction for Historical Documents