Resources for Turkish Natural Language Processing: A critical survey
arXiv:2204.05042 · doi:10.1007/s10579-022-09605-4
Abstract
This paper presents a comprehensive survey of corpora and lexical resources available for Turkish. We review a broad range of resources, focusing on the ones that are publicly available. In addition to providing information about the available linguistic resources, we present a set of recommendations, and identify gaps in the data available for conducting research and building applications in Turkish Linguistics and Natural Language Processing.
Published in Language Resources and Evaluation
References in corpus (6)
- SemEval-2017 Task 1: Semantic Textual Similarity - Multilingual and Cross-lingual Focused Evaluation
- AUTSL: A Large Scale Multi-modal Turkish Sign Language Dataset and Baseline Methods
- GeoCoV19: A Dataset of Hundreds of Millions of Multilingual COVID-19 Tweets with Location Information
- Critical Survey of the Freely Available Arabic Corpora
- A Monolingual Approach to Contextualized Word Embeddings for Mid-Resource Languages
- MediaSpeech: Multilanguage ASR Benchmark and Dataset