activity
20182024
most citedTranslation Transformers Rediscover Inherent Data Domains

5 citations · 6 across the 3 of their papers we have counts for

collaborators

5 papers

cs.CL2024

BPE Gets Picky: Efficient Vocabulary Refinement During Tokenizer Training

Pavel Chizhov, Catherine Arnett, Elizaveta Korotkova +1

Language models can largely benefit from efficient tokenization. However, they still mostly utilize the classical BPE algorithm, a simple and reliable method. This has been shown t…

cs.CL2023★ 1 cited

Beyond Toxic: Toxicity Detection Datasets are Not Enough for Brand Safety

Elizaveta Korotkova, Isaac Chung

The rapid growth in user generated content on social media has resulted in a significant rise in demand for automated content moderation. Various methods and frameworks have been p…

cs.CL2021★ 5 cited

Translation Transformers Rediscover Inherent Data Domains

Maksym Del, Elizaveta Korotkova, Mark Fishel

Many works proposed methods to improve the performance of Neural Machine Translation (NMT) models in a domain/multi-domain adaptation scenario. However, an understanding of how NMT…

cs.CL2019

Grammatical Error Correction and Style Transfer via Zero-shot Monolingual Translation

Elizaveta Korotkova, Agnes Luhtaru, Maksym Del +3

Both grammatical error correction and text style transfer can be viewed as monolingual sequence-to-sequence transformation tasks, but the scarcity of directly annotated data for ei…

cs.CL2018

Monolingual and Cross-lingual Zero-shot Style Transfer

Elizaveta Korotkova, Maksym Del, Mark Fishel

We introduce the task of zero-shot style transfer between different languages. Our training data includes multilingual parallel corpora, but does not contain any parallel sentences…