activity
20212023
most citedEDS-MEMBED: Multi-sense embeddings based on enhanced distributional semantic structures via a graph walk over word senses

13 citations · 17 across the 5 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2023

People and Places of Historical Europe: Bootstrapping Annotation Pipeline and a New Corpus of Named Entities in Late Medieval Texts

Vít Novotný, Kristýna Luger, Michal Štefánik +2

Although pre-trained named entity recognition (NER) models are highly accurate on modern corpora, they underperform on historical texts due to differences in language OCR errors. I…

cs.CL2022

Adaptor: Objective-Centric Adaptation Framework for Language Models

Michal Štefánik, Vít Novotný, Nikola Groverová +1

Progress in natural language processing research is catalyzed by the possibilities given by the widespread software frameworks. This paper introduces Adaptor library that transpose…

cs.CL2021★ 4 cited

Regressive Ensemble for Machine Translation Quality Evaluation

Michal Štefánik, Vít Novotný, Petr Sojka

This work introduces a simple regressive ensemble for evaluating machine translation quality based on a set of novel and established metrics. We evaluate the ensemble using a corre…

cs.CL2021★ 13 cited

EDS-MEMBED: Multi-sense embeddings based on enhanced distributional semantic structures via a graph walk over word senses

Eniafe Festus Ayetiran, Petr Sojka, Vít Novotný

Several language applications often require word semantics as a core part of their processing pipeline, either as precise meaning inference or semantic similarity. Multi-sense embe…

cs.CL2021

One Size Does Not Fit All: Finding the Optimal Subword Sizes for FastText Models across Languages

Vít Novotný, Eniafe Festus Ayetiran, Dalibor Bačovský +3

Unsupervised representation learning of words from large multilingual corpora is useful for downstream tasks such as word sense disambiguation, semantic text similarity, and inform…