6 papers
The GINCO Training Dataset for Web Genre Identification of Documents Out in the Wild
Taja Kuzman, Peter Rupnik, Nikola Ljubešić
This paper presents a new training dataset for automatic genre identification GINCO, which is based on 1,125 crawled Slovenian web documents that consist of 650 thousand words. Eac…
Community evolution in retweet networks
Bojan Evkoski, Igor Mozetic, Nikola Ljubesic +1
Communities in social networks often reflect close social ties between their members and their evolution through time. We propose an approach that tracks two aspects of community e…
CoSimLex: A Resource for Evaluating Graded Word Similarity in Context
Carlos Santos Armendariz, Matthew Purver, Matej Ulčar +5
State of the art natural language processing tools are built on context-dependent word embeddings, but no direct method for evaluating these representations currently exists. Stand…
The FRENK Datasets of Socially Unacceptable Discourse in Slovene and English
Nikola Ljubešić, Darja Fišer, Tomaž Erjavec
In this paper we present datasets of Facebook comment threads to mainstream media posts in Slovene and English developed inside the Slovene national project FRENK which cover two t…
KAS-term: Extracting Slovene Terms from Doctoral Theses via Supervised Machine Learning
Nikola Ljubešić, Darja Fišer, Tomaž Erjavec
This paper presents a dataset and supervised learning experiments for term extraction from Slovene academic texts. Term candidates in the dataset were extracted via morphosyntactic…
Bleaching Text: Abstract Features for Cross-lingual Gender Prediction
Rob van der Goot, Nikola Ljubešić, Ian Matroos +2
Gender prediction has typically focused on lexical and social network features, yielding good performance, but making systems highly language-, topic-, and platform-dependent. Cros…