4 papers
SiNFluD: Creating and Evaluating Figurative Language Dataset for Sindhi
Wazir Ali, Adeeb Noor, Saifullah Tumrani
In this article, we introduce SiNFluD, a novel benchmark dataset for Sindhi figurative language classification. We first collect raw text from various blogs, social media platforms…
BioUNER: A Benchmark Dataset for Clinical Urdu Named Entity Recognition
Wazir Ali, Adeeb Noor, Sanaullah Mahar +2
In this article, we present a gold-standard benchmark dataset for Biomedical Urdu Named Entity Recognition (BioUNER), developed by crawling health-related articles from online Urdu…
Enhancing Sindhi Word Segmentation using Subword Representation Learning and Position-aware Self-attention
Wazir Ali, Jay Kumar, Saifullah Tumrani +3
Sindhi word segmentation is a challenging task due to space omission and insertion issues. The Sindhi language itself adds to this complexity. It's cursive and consists of characte…
An Evaluation of Sindhi Word Embedding in Semantic Analogies and Downstream Tasks
Wazir Ali, Saifullah Tumrani, Jay Kumar +1
In this paper, we propose a new word embedding based corpus consisting of more than 61 million words crawled from multiple web resources. We design a preprocessing pipeline for the…