paper

NADI 2020: The First Nuanced Arabic Dialect Identification Shared Task

arXiv:2010.11334

Abstract

We present the results and findings of the First Nuanced Arabic Dialect Identification Shared Task (NADI). This Shared Task includes two subtasks: country-level dialect identification (Subtask 1) and province-level sub-dialect identification (Subtask 2). The data for the shared task covers a total of 100 provinces from 21 Arab countries and are collected from the Twitter domain. As such, NADI is the first shared task to target naturally-occurring fine-grained dialectal text at the sub-country level. A total of 61 teams from 25 countries registered to participate in the tasks, thus reflecting the interest of the community in this area. We received 47 submissions for Subtask 1 from 18 teams and 9 submissions for Subtask 2 from 9 teams.

Accepted in The Fifth Arabic Natural Language Processing Workshop (WANLP 2020)

References in corpus (3)

NADI 2020: The First Nuanced Arabic Dialect Identification Shared Task · wovepaper