Publications (11)
Beyond Plain Toxic: Detection of Inappropriate Statements on Flammable Topics for the Russian Language
Nikolay Babakov, Varvara Logacheva, Alexander Panchenko
Toxicity on the Internet, such as hate speech, offenses towards particular users or groups of people, or the use of obscene words, is an acknowledged problem. However, there also e…
Word Sense Disambiguation for 158 Languages using Word Embeddings Only
Varvara Logacheva, Denis Teslenko, Artem Shelmanov +7
Disambiguation of word senses in context is easy for humans, but is a major challenge for automatic approaches. Sophisticated supervised and knowledge-based models were developed t…
Text Detoxification using Large Pre-trained Neural Models
David Dale, Anton Voronov, Daryna Dementieva +4
We present two novel unsupervised methods for eliminating toxicity in text. Our first method combines two recent ideas: (1) guidance of the generation process with small style-cond…
RUSSE'2020: Findings of the First Taxonomy Enrichment Task for the Russian language
Irina Nikishina, Varvara Logacheva, Alexander Panchenko +1
This paper describes the results of the first shared task on taxonomy enrichment for the Russian language. The participants were asked to extend an existing taxonomy with previousl…
Studying the role of named entities for content preservation in text style transfer
Nikolay Babakov, David Dale, Varvara Logacheva +2
Text style transfer techniques are gaining popularity in Natural Language Processing, finding various applications such as text detoxification, sentiment, or formality transfer. Ho…
Studying Taxonomy Enrichment on Diachronic WordNet Versions
Irina Nikishina, Alexander Panchenko, Varvara Logacheva +1
Ontologies, taxonomies, and thesauri are used in many NLP tasks. However, most studies are focused on the creation of these lexical resources rather than the maintenance of the exi…
Few-shot classification in Named Entity Recognition Task
Alexander Fritzler, Varvara Logacheva, Maksim Kretov
For many natural language processing (NLP) tasks the amount of annotated data is limited. This urges a need to apply semi-supervised learning techniques, such as transfer learning…
The Second Conversational Intelligence Challenge (ConvAI2)
Emily Dinan, Varvara Logacheva, Valentin Malykh +14
We describe the setting and results of the ConvAI2 NeurIPS competition that aims to further the state-of-the-art in open-domain chatbots. Some key takeaways from the competition ar…
Taxonomy Enrichment with Text and Graph Vector Representations
Irina Nikishina, Mikhail Tikhomirov, Varvara Logacheva +3
Knowledge graphs such as DBpedia, Freebase or Wikidata always contain a taxonomic backbone that allows the arrangement and structuring of various concepts in accordance with the hy…
Methods for Detoxification of Texts for the Russian Language
Daryna Dementieva, Daniil Moskovskiy, Varvara Logacheva +4
We introduce the first study of automatic detoxification of Russian texts to combat offensive language. Such a kind of textual style transfer can be used, for instance, for process…
Detecting Inappropriate Messages on Sensitive Topics that Could Harm a Company's Reputation
Nikolay Babakov, Varvara Logacheva, Olga Kozlova +2
Not all topics are equally "flammable" in terms of toxicity: a calm discussion of turtles or fishing less often fuels inappropriate toxic dialogues than a discussion of politics or…