paper

A Multi-Stage Framework for Kuzushiji Character Recognition in Japanese Historical Documents

arXiv:2602.19086

Abstract

Kuzushiji was a widely used cursive writing system in pre-modern Japan. Due to simplification and glyph variation, most modern Japanese readers cannot read Kuzushiji characters. Consequently, recent studies have developed optical character recognition (OCR) systems for Kuzushiji. Despite recent progress, Kuzushiji character recognition (KCR) in Japanese historical documents remains challenging because of seal-character overlap and complex layouts, which interfere with character recognition and hinder accurate reconstruction of the reading order. To address these challenges, we propose a multi-stage KCR framework comprising character detection, cropping, classification, ordering, and large language model (LLM)-based post-OCR correction. Specifically, we employ a synthetic data augmentation strategy to improve character detection robustness against seal interference and introduce an adaptive column clustering algorithm to reconstruct the reading order. Finally, we leverage the contextual capabilities of the LLM to correct OCR errors. In addition, we correct annotation omissions, reconstruct the benchmark dataset, and introduce a synthetic test set with simulated seal interference and an out-of-domain (OOD) test set for evaluation. Compared with the conventional character-level OCR baseline, our framework achieves relative CER reductions of 43.48%, 46.02%, and 39.11% on the real, synthetic, and OOD test sets, respectively.

Project page is available at https://ruiyangju.github.io/KuzushijiOCR/