1 citations · 1 across the 4 of their papers we have counts for
Showing 2024Show all
2 papers · 1 filter
cs.CL2024★ 1 cited
Scrambled text: training Language Models to correct OCR errors using synthetic data
Jonathan Bourne
OCR errors are common in digitised historical archives significantly affecting their usability and value. Generative Language Models (LMs) have shown potential for correcting these…
cs.CL2024
CLOCR-C: Context Leveraging OCR Correction with Pre-trained Language Models
Jonathan Bourne
The digitisation of historical print media archives is crucial for increasing accessibility to contemporary records. However, the process of Optical Character Recognition (OCR) use…