3 papers
cs.CV2026
Koshur Pixel: a large-scale synthetic ocr dataset for kashmiri
Haq Nawaz Malik, Faizan Iqbal, Nahfid Nissar
Optical Character Recognition (OCR) for low-resource languages is often constrained by the lack of annotated training data and the complexity of script-specific rendering. Kashmiri…
cs.CL2026
Koshur Diacritizer: A Byte-Level Sequence-to-Sequence Model for Kashmiri Diacritic Restoration
Haq Nawaz Malik, Nahfid Nissar, Faizan Iqbal
Kashmiri, an Indo-Aryan language written in a modified Perso-Arabic script, frequently omits diacritic marks in digital text, creating ambiguity and challenging downstream NLP appl…
cs.CL2026
ks-pret-5m: a 5 million word, 12 million token kashmiri pretraining dataset
Haq Nawaz Malik, Nahfid Nissar
We present KS-PRET-5M, the largest publicly available pretraining dataset for the Kashmiri language, comprising 5,090,244 (5.09M) words, 27,692,959 (27.6M) characters, and a vocabu…