Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
Reading the unreadable: Creating a dataset of 19th century English newspapers using image-to-text language models
Jonathan Bourne
Oscar Wilde said, "The difference between literature and journalism is that journalism is unreadable, and literature is not read." Unfortunately, The digitally archived journalism…
cs.CL2025
CLOCR-C: Context Leveraging OCR Correction with Pre-trained Language Models
Jonathan Bourne
The digitisation of historical print media archives is crucial for increasing accessibility to contemporary records. However, the process of Optical Character Recognition (OCR) use…
cs.CL2024
Scrambled text: training Language Models to correct OCR errors using synthetic data
Jonathan Bourne
OCR errors are common in digitised historical archives significantly affecting their usability and value. Generative Language Models (LMs) have shown potential for correcting these…