37 citations · 38 across the 6 of their papers we have counts for
9 papers · 1 filter
Developing Universal Dependency Treebanks for Magahi and Braj
Mohit Raj, Shyam Ratan, Deepak Alok +2
In this paper, we discuss the development of treebanks for two low-resourced Indian languages - Magahi and Braj based on the Universal Dependencies framework. The Magahi treebank c…
Language Resources and Technologies for Non-Scheduled and Endangered Indian Languages
Ritesh Kumar, Bornini Lahiri
In the present paper, we will present a survey of the language resources and technologies available for the non-scheduled and endangered languages of India. While there have been d…
Demo of the Linguistic Field Data Management and Analysis System -- LiFE
Siddharth Singh, Ritesh Kumar, Shyam Ratan +1
In the proposed demo, we will present a new software - Linguistic Field Data Management and Analysis System - LiFE (https://github.com/kmi-linguistics/life) - an open-source, web-b…
SIGTYP 2021 Shared Task: Robust Spoken Language Identification
Elizabeth Salesky, Badr M. Abdullah, Sabrina J. Mielke +6
While language identification is a fundamental speech and language processing task, for many languages and language families it remains a challenging task. For many low-resource an…
Developing a Multilingual Annotated Corpus of Misogyny and Aggression
Shiladitya Bhattacharya, Siddharth Singh, Ritesh Kumar +5
In this paper, we discuss the development of a multilingual annotated corpus of misogyny and aggression in Indian English, Hindi, and Indian Bangla as part of a project on studying…
SemEval-2019 Task 6: Identifying and Categorizing Offensive Language in Social Media (OffensEval)
Marcos Zampieri, Shervin Malmasi, Preslav Nakov +3
We present the results and the main findings of SemEval-2019 Task 6 on Identifying and Categorizing Offensive Language in Social Media (OffensEval). The task was based on a new dat…