7 papers
Static Embeddings as Efficient Knowledge Bases?
Philipp Dufter, Nora Kassner, Hinrich Schütze
Recent research investigates factual knowledge stored in large pretrained language models (PLMs). Instead of structural knowledge base (KB) queries, masked sentences such as "Paris…
Multilingual LAMA: Investigating Knowledge in Multilingual Pretrained Language Models
Nora Kassner, Philipp Dufter, Hinrich Schütze
Recently, it has been found that monolingual English language models can be used as knowledge bases. Instead of structural knowledge base queries, masked sentences such as "Paris i…
Identifying Necessary Elements for BERT's Multilinguality
Philipp Dufter, Hinrich Schütze
It has been shown that multilingual BERT (mBERT) yields high quality multilingual representations and enables effective zero-shot transfer. This is surprising given that mBERT does…
Quantifying the Contextualization of Word Representations with Semantic Class Probing
Mengjie Zhao, Philipp Dufter, Yadollah Yaghoobzadeh +1
Pretrained language models have achieved a new state of the art on many NLP tasks, but there are still many open questions about how and why they work so well. We investigate the c…
SimAlign: High Quality Word Alignments without Parallel Training Data using Static and Contextualized Embeddings
Masoud Jalili Sabet, Philipp Dufter, François Yvon +1
Word alignments are useful for tasks like statistical and neural machine translation (NMT) and cross-lingual annotation projection. Statistical word aligners perform well, as do me…
Analytical Methods for Interpretable Ultradense Word Embeddings
Philipp Dufter, Hinrich Schütze
Word embeddings are useful for a wide variety of tasks, but they lack interpretability. By rotating word spaces, interpretable dimensions can be identified while preserving the inf…