activity
20212023
most citedEDS-MEMBED: Multi-sense embeddings based on enhanced distributional semantic structures via a graph walk over word senses

13 citations · 17 across the 5 of their papers we have counts for

collaborators

6 papers

cs.CL2023

People and Places of Historical Europe: Bootstrapping Annotation Pipeline and a New Corpus of Named Entities in Late Medieval Texts

Vít Novotný, Kristýna Luger, Michal Štefánik +2

Although pre-trained named entity recognition (NER) models are highly accurate on modern corpora, they underperform on historical texts due to differences in language OCR errors. I…

cs.CL2022

Adaptor: Objective-Centric Adaptation Framework for Language Models

Michal Štefánik, Vít Novotný, Nikola Groverová +1

Progress in natural language processing research is catalyzed by the possibilities given by the widespread software frameworks. This paper introduces Adaptor library that transpose…

cs.CL20214 cited

Regressive Ensemble for Machine Translation Quality Evaluation

Michal Štefánik, Vít Novotný, Petr Sojka

This work introduces a simple regressive ensemble for evaluating machine translation quality based on a set of novel and established metrics. We evaluate the ensemble using a corre…

cs.DL2021

WebMIaS on Docker: Deploying Math-Aware Search in a Single Line of Code

Dávid Lupták, Vít Novotný, Michal Štefánik +1

Math informational retrieval (MIR) search engines are absent in the wide-spread production use, even though documents in the STEM fields contain many mathematical formulae, which a…

cs.CL202113 cited

EDS-MEMBED: Multi-sense embeddings based on enhanced distributional semantic structures via a graph walk over word senses

Eniafe Festus Ayetiran, Petr Sojka, Vít Novotný

Several language applications often require word semantics as a core part of their processing pipeline, either as precise meaning inference or semantic similarity. Multi-sense embe…

cs.CL2021

One Size Does Not Fit All: Finding the Optimal Subword Sizes for FastText Models across Languages

Vít Novotný, Eniafe Festus Ayetiran, Dalibor Bačovský +3

Unsupervised representation learning of words from large multilingual corpora is useful for downstream tasks such as word sense disambiguation, semantic text similarity, and inform…