activity
20152021
most citedHow Language-Neutral is Multilingual BERT?

76 citations · 91 across the 3 of their papers we have counts for

collaborators
Showing cs.CLShow all

9 papers · 1 filter

cs.CL2021

THEaiTRE 1.0: Interactive generation of theatre play scripts

Rudolf Rosa, Tomáš Musil, Ondřej Dušek +13

We present the first version of a system for interactive generation of theatre play scripts. The system is based on a vanilla GPT-2 model with several adjustments, targeting specif…

cs.CL2020

Predicting Typological Features in WALS using Language Embeddings and Conditional Probabilities: ÚFAL Submission to the SIGTYP 2020 Shared Task

Martin Vastl, Daniel Zeman, Rudolf Rosa

We present our submission to the SIGTYP 2020 Shared Task on the prediction of typological features. We submit a constrained system, predicting typological features only based on th…

cs.CL2020

Universal Dependencies according to BERT: both more specific and more general

Tomasz Limisiewicz, Rudolf Rosa, David Mareček

This work focuses on analyzing the form and extent of syntactic abstraction captured by BERT by extracting labeled dependency trees from self-attentions. Previous work showed that…

cs.CL2020

On the Language Neutrality of Pre-trained Multilingual Representations

Jindřich Libovický, Rudolf Rosa, Alexander Fraser

Multilingual contextual embeddings, such as multilingual BERT and XLM-RoBERTa, have proved useful for many multi-lingual tasks. Previous work probed the cross-linguality of the rep…

cs.CL201976 cited

How Language-Neutral is Multilingual BERT?

Jindřich Libovický, Rudolf Rosa, Alexander Fraser

Multilingual BERT (mBERT) provides sentence representations for 104 languages, which are useful for many multi-lingual tasks. Previous work probed the cross-linguality of mBERT usi…

cs.CL2019

Unsupervised Lemmatization as Embeddings-Based Word Clustering

Rudolf Rosa, Zdeněk Žabokrtský

We focus on the task of unsupervised lemmatization, i.e. grouping together inflected forms of one word under one label (a lemma) without the use of annotated training data. We prop…