2 citations · 2 across the 3 of their papers we have counts for
6 papers · 1 filter
CroSentiNews 2.0: A Sentence-Level News Sentiment Corpus
Gaurish Thakkar, Nives Mikelic Preradović, Marko Tadić
This article presents a sentence-level sentiment dataset for the Croatian news domain. In addition to the 3K annotated texts already present, our dataset contains 14.5K annotated s…
Croatian Film Review Dataset (Cro-FiReDa): A Sentiment Annotated Dataset of Film Reviews
Gaurish Thakkar, Nives Mikelic Preradovic, Marko Tadić
This paper introduces Cro-FiReDa, a sentiment-annotated dataset for Croatian in the domain of movie reviews. The dataset, which contains over 10,000 sentences, has been annotated a…
Natural Language Processing Chains Inside a Cross-lingual Event-Centric Knowledge Pipeline for European Union Under-resourced Languages
Diego Alves, Gaurish Thakkar, Marko Tadić
This article presents the strategy for developing a platform containing Language Processing Chains for European Union languages, consisting of Tokenization to Parsing, also includi…
Evaluating Language Tools for Fifteen EU-official Under-resourced Languages
Diego Alves, Gaurish Thakkar, Marko Tadić
This article presents the results of the evaluation campaign of language tools available for fifteen EU-official under-resourced languages. The evaluation was conducted within the…
UNER: Universal Named-Entity RecognitionFramework
Diego Alves, Tin Kuculo, Gabriel Amaral +2
We introduce the Universal Named-Entity Recognition (UNER)framework, a 4-level classification hierarchy, and the methodology that isbeing adopted to create the first multilingual U…
The European Language Technology Landscape in 2020: Language-Centric and Human-Centric AI for Cross-Cultural Communication in Multilingual Europe
Georg Rehm, Katrin Marheinecke, Stefanie Hegele +44
Multilingualism is a cultural cornerstone of Europe and firmly anchored in the European treaties including full language equality. However, language barriers impacting business, cr…