activity
20182022
collaborators

6 papers

cs.CL2022

The GINCO Training Dataset for Web Genre Identification of Documents Out in the Wild

Taja Kuzman, Peter Rupnik, Nikola Ljubešić

This paper presents a new training dataset for automatic genre identification GINCO, which is based on 1,125 crawled Slovenian web documents that consist of 650 thousand words. Eac…

cs.SI2021

Community evolution in retweet networks

Bojan Evkoski, Igor Mozetic, Nikola Ljubesic +1

Communities in social networks often reflect close social ties between their members and their evolution through time. We propose an approach that tracks two aspects of community e…

cs.CL2019

CoSimLex: A Resource for Evaluating Graded Word Similarity in Context

Carlos Santos Armendariz, Matthew Purver, Matej Ulčar +5

State of the art natural language processing tools are built on context-dependent word embeddings, but no direct method for evaluating these representations currently exists. Stand…

cs.CL2019

The FRENK Datasets of Socially Unacceptable Discourse in Slovene and English

Nikola Ljubešić, Darja Fišer, Tomaž Erjavec

In this paper we present datasets of Facebook comment threads to mainstream media posts in Slovene and English developed inside the Slovene national project FRENK which cover two t…

cs.CL2019

KAS-term: Extracting Slovene Terms from Doctoral Theses via Supervised Machine Learning

Nikola Ljubešić, Darja Fišer, Tomaž Erjavec

This paper presents a dataset and supervised learning experiments for term extraction from Slovene academic texts. Term candidates in the dataset were extracted via morphosyntactic…

cs.CL2018

Bleaching Text: Abstract Features for Cross-lingual Gender Prediction

Rob van der Goot, Nikola Ljubešić, Ian Matroos +2

Gender prediction has typically focused on lexical and social network features, yielding good performance, but making systems highly language-, topic-, and platform-dependent. Cros…