7 papers
Persian Pixel: A large-scale synthetic OCR dataset for Persian language
Pouria Mahdi, Haq Nawaz Malik
Optical Character Recognition (OCR) for Persian remains substantially less mature than for Latin-script languages despite Persian being spoken by more than 110 million people acros…
Koshur Pixel: a large-scale synthetic ocr dataset for kashmiri
Haq Nawaz Malik, Faizan Iqbal, Nahfid Nissar
Optical Character Recognition (OCR) for low-resource languages is often constrained by the lack of annotated training data and the complexity of script-specific rendering. Kashmiri…
Koshur Diacritizer: A Byte-Level Sequence-to-Sequence Model for Kashmiri Diacritic Restoration
Haq Nawaz Malik, Nahfid Nissar, Faizan Iqbal
Kashmiri, an Indo-Aryan language written in a modified Perso-Arabic script, frequently omits diacritic marks in digital text, creating ambiguity and challenging downstream NLP appl…
ks-pret-5m: a 5 million word, 12 million token kashmiri pretraining dataset
Haq Nawaz Malik, Nahfid Nissar
We present KS-PRET-5M, the largest publicly available pretraining dataset for the Kashmiri language, comprising 5,090,244 (5.09M) words, 27,692,959 (27.6M) characters, and a vocabu…
synthocr-gen: A synthetic ocr dataset generator for low-resource languages- breaking the data barrier
Haq Nawaz Malik, Kh Mohmad Shafi, Tanveer Ahmad Reshi
Optical Character Recognition (OCR) for low-resource languages remains a significant challenge due to the scarcity of large-scale annotated training datasets. Languages such as Kas…
ks-lit-3m: A 3.1 million word kashmiri text dataset for large language model pretraining
Haq Nawaz Malik
Large Language Models (LLMs) demonstrate remarkable fluency across high-resource languages yet consistently fail to generate coherent text in Kashmiri, a language spoken by approxi…