2 papers
cs.CL2026
SiNFluD: Creating and Evaluating Figurative Language Dataset for Sindhi
Wazir Ali, Adeeb Noor, Saifullah Tumrani
In this article, we introduce SiNFluD, a novel benchmark dataset for Sindhi figurative language classification. We first collect raw text from various blogs, social media platforms…
cs.CL2026
BioUNER: A Benchmark Dataset for Clinical Urdu Named Entity Recognition
Wazir Ali, Adeeb Noor, Sanaullah Mahar +2
In this article, we present a gold-standard benchmark dataset for Biomedical Urdu Named Entity Recognition (BioUNER), developed by crawling health-related articles from online Urdu…