activity
20192024
most citedCross-lingual alignments of ELMo contextual embeddings

24 citations · 50 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CL2024★ 5 cited

Code-mixed Sentiment and Hate-speech Prediction

Anjali Yadav, Tanya Garg, Matej Klemen +3

Code-mixed discourse combines multiple languages in a single text. It is commonly used in informal discourse in countries with several official languages, but also in many other co…

cs.CL2022★ 3 cited

Sequence to sequence pretraining for a less-resourced Slovenian language

Matej Ulčar, Marko Robnik-Šikonja

Large pretrained language models have recently conquered the area of natural language processing. As an alternative to predominant masked language modelling introduced in BERT, the…

cs.CL2021★ 24 cited

Cross-lingual alignments of ELMo contextual embeddings

Matej Ulčar, Marko Robnik-Šikonja

Building machine learning prediction models for a specific NLP task requires sufficient training data, which can be difficult to obtain for less-resourced languages. Cross-lingual…

cs.CL2019★ 14 cited

CoSimLex: A Resource for Evaluating Graded Word Similarity in Context

Carlos Santos Armendariz, Matthew Purver, Matej Ulčar +5

State of the art natural language processing tools are built on context-dependent word embeddings, but no direct method for evaluating these representations currently exists. Stand…

cs.CL2019★ 4 cited

Multilingual Culture-Independent Word Analogy Datasets

Matej Ulčar, Kristiina Vaik, Jessica Lindström +2

In text processing, deep neural networks mostly use word embeddings as an input. Embeddings have to ensure that relations between words are reflected through distances in a high-di…