activity
20182021
collaborators

7 papers

cs.CL2021

Static Embeddings as Efficient Knowledge Bases?

Philipp Dufter, Nora Kassner, Hinrich Schütze

Recent research investigates factual knowledge stored in large pretrained language models (PLMs). Instead of structural knowledge base (KB) queries, masked sentences such as "Paris…

cs.CL2021

Multilingual LAMA: Investigating Knowledge in Multilingual Pretrained Language Models

Nora Kassner, Philipp Dufter, Hinrich Schütze

Recently, it has been found that monolingual English language models can be used as knowledge bases. Instead of structural knowledge base queries, masked sentences such as "Paris i…

cs.CL2020

Identifying Necessary Elements for BERT's Multilinguality

Philipp Dufter, Hinrich Schütze

It has been shown that multilingual BERT (mBERT) yields high quality multilingual representations and enables effective zero-shot transfer. This is surprising given that mBERT does…

cs.CL2020

Quantifying the Contextualization of Word Representations with Semantic Class Probing

Mengjie Zhao, Philipp Dufter, Yadollah Yaghoobzadeh +1

Pretrained language models have achieved a new state of the art on many NLP tasks, but there are still many open questions about how and why they work so well. We investigate the c…

cs.CL2020

SimAlign: High Quality Word Alignments without Parallel Training Data using Static and Contextualized Embeddings

Masoud Jalili Sabet, Philipp Dufter, François Yvon +1

Word alignments are useful for tasks like statistical and neural machine translation (NMT) and cross-lingual annotation projection. Statistical word aligners perform well, as do me…

cs.CL2019

Analytical Methods for Interpretable Ultradense Word Embeddings

Philipp Dufter, Hinrich Schütze

Word embeddings are useful for a wide variety of tasks, but they lack interpretability. By rotating word spaces, interpretable dimensions can be identified while preserving the inf…