Showing 2021Show all
3 papers · 1 filter
cs.CL2021
Language Identification of Hindi-English tweets using code-mixed BERT
Mohd Zeeshan Ansari, M M Sufyan Beg, Tanvir Ahmad +2
Language identification of social media text has been an interesting problem of study in recent years. Social media messages are predominantly in code mixed in non-English speaking…
cs.CL2021
Language Lexicons for Hindi-English Multilingual Text Processing
Mohd Zeeshan Ansari, Tanvir Ahmad, Noaima Bari
Language Identification in textual documents is the process of automatically detecting the language contained in a document based on its content. The present Language Identificatio…
cs.CL2021
A Simple and Efficient Probabilistic Language model for Code-Mixed Text
M Zeeshan Ansari, Tanvir Ahmad, M M Sufyan Beg +1
The conventional natural language processing approaches are not accustomed to the social media text due to colloquial discourse and non-homogeneous characteristics. Significantly,…