collaborators

7 papers

cs.CV2026

Persian Pixel: A large-scale synthetic OCR dataset for Persian language

Pouria Mahdi, Haq Nawaz Malik

Optical Character Recognition (OCR) for Persian remains substantially less mature than for Latin-script languages despite Persian being spoken by more than 110 million people acros…

cs.CV2026

Koshur Pixel: a large-scale synthetic ocr dataset for kashmiri

Haq Nawaz Malik, Faizan Iqbal, Nahfid Nissar

Optical Character Recognition (OCR) for low-resource languages is often constrained by the lack of annotated training data and the complexity of script-specific rendering. Kashmiri…

cs.CL2026

Koshur Diacritizer: A Byte-Level Sequence-to-Sequence Model for Kashmiri Diacritic Restoration

Haq Nawaz Malik, Nahfid Nissar, Faizan Iqbal

Kashmiri, an Indo-Aryan language written in a modified Perso-Arabic script, frequently omits diacritic marks in digital text, creating ambiguity and challenging downstream NLP appl…

cs.CL2026

ks-pret-5m: a 5 million word, 12 million token kashmiri pretraining dataset

Haq Nawaz Malik, Nahfid Nissar

We present KS-PRET-5M, the largest publicly available pretraining dataset for the Kashmiri language, comprising 5,090,244 (5.09M) words, 27,692,959 (27.6M) characters, and a vocabu…

cs.CL2026

synthocr-gen: A synthetic ocr dataset generator for low-resource languages- breaking the data barrier

Haq Nawaz Malik, Kh Mohmad Shafi, Tanveer Ahmad Reshi

Optical Character Recognition (OCR) for low-resource languages remains a significant challenge due to the scarcity of large-scale annotated training datasets. Languages such as Kas…

cs.CL2026

ks-lit-3m: A 3.1 million word kashmiri text dataset for large language model pretraining

Haq Nawaz Malik

Large Language Models (LLMs) demonstrate remarkable fluency across high-resource languages yet consistently fail to generate coherent text in Kashmiri, a language spoken by approxi…