2 citations · 4 across the 4 of their papers we have counts for
4 papers
Creating and Managing a large annotated parallel corpora of Indian languages
Ritesh Kumar, Shiv Bhusan Kaushik, Pinkey Nainwani +1
This paper presents the challenges in creating and managing large parallel corpora of 12 major Indian languages (which is soon to be extended to 23 languages) as part of a major co…
Challenges in Developing LRs for Non-Scheduled Languages: A Case of Magahi
Ritesh Kumar
Magahi is an Indo-Aryan Language, spoken mainly in the Eastern parts of India. Despite having a significant number of speakers, there has been virtually no language resource (LR) o…
Towards automatic identification of linguistic politeness in Hindi texts
Ritesh Kumar
In this paper I present a classifier for automatic identification of linguistic politeness in Hindi texts. I have used the manually annotated corpus of over 25,000 blog comments to…
The ComMA Dataset V0.2: Annotating Aggression and Bias in Multilingual Social Media Discourse
Ritesh Kumar, Enakshi Nandi, Laishram Niranjana Devi +4
In this paper, we discuss the development of a multilingual dataset annotated with a hierarchical, fine-grained tagset marking different types of aggression and the "context" in wh…