24 citations · 50 across the 5 of their papers we have counts for
5 papers
Code-mixed Sentiment and Hate-speech Prediction
Anjali Yadav, Tanya Garg, Matej Klemen +3
Code-mixed discourse combines multiple languages in a single text. It is commonly used in informal discourse in countries with several official languages, but also in many other co…
Sequence to sequence pretraining for a less-resourced Slovenian language
Matej Ulčar, Marko Robnik-Šikonja
Large pretrained language models have recently conquered the area of natural language processing. As an alternative to predominant masked language modelling introduced in BERT, the…
Cross-lingual alignments of ELMo contextual embeddings
Matej Ulčar, Marko Robnik-Šikonja
Building machine learning prediction models for a specific NLP task requires sufficient training data, which can be difficult to obtain for less-resourced languages. Cross-lingual…
CoSimLex: A Resource for Evaluating Graded Word Similarity in Context
Carlos Santos Armendariz, Matthew Purver, Matej Ulčar +5
State of the art natural language processing tools are built on context-dependent word embeddings, but no direct method for evaluating these representations currently exists. Stand…
Multilingual Culture-Independent Word Analogy Datasets
Matej Ulčar, Kristiina Vaik, Jessica Lindström +2
In text processing, deep neural networks mostly use word embeddings as an input. Embeddings have to ensure that relations between words are reflected through distances in a high-di…