most citedInvestigating Machine Learning Methods for Language and Dialect Identification of Cuneiform Texts

9 citations · 31 across the 6 of their papers we have counts for

collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL20217 cited

SINA-BERT: A pre-trained Language Model for Analysis of Medical Texts in Persian

Nasrin Taghizadeh, Ehsan Doostmohammadi, Elham Seifossadat +2

We have released Sina-BERT, a language model pre-trained on BERT (Devlin et al., 2018) to address the lack of a high-quality Persian language model in the medical domain. SINA-BERT…

cs.CL2020

Joint Persian Word Segmentation Correction and Zero-Width Non-Joiner Recognition Using BERT

Ehsan Doostmohammadi, Minoo Nassajian, Adel Rahimi

Words are properly segmented in the Persian writing system; in practice, however, these writing rules are often neglected, resulting in single words being written disjointedly and…

cs.CL2020

Persian Ezafe Recognition Using Transformers and Its Role in Part-Of-Speech Tagging

Ehsan Doostmohammadi, Minoo Nassajian, Adel Rahimi

Ezafe is a grammatical particle in some Iranian languages that links two words together. Regardless of the important information it conveys, it is almost always not indicated in Pe…

cs.CL20203 cited

Persian Keyphrase Generation Using Sequence-to-Sequence Models

Ehsan Doostmohammadi, Mohammad Hadi Bokaei, Hossein Sameti

Keyphrases are a very short summary of an input text and provide the main subjects discussed in the text. Keyphrase extraction is a useful upstream task and can be used in various…

cs.CL20204 cited

PerKey: A Persian News Corpus for Keyphrase Extraction and Generation

Ehsan Doostmohammadi, Mohammad Hadi Bokaei, Hossein Sameti

Keyphrases provide an extremely dense summary of a text. Such information can be used in many Natural Language Processing tasks, such as information retrieval and text summarizatio…

cs.CL20209 cited

Investigating Machine Learning Methods for Language and Dialect Identification of Cuneiform Texts

Ehsan Doostmohammadi, Minoo Nassajian

Identification of the languages written using cuneiform symbols is a difficult task due to the lack of resources and the problem of tokenization. The Cuneiform Language Identificat…