activity
20182021
collaborators
Showing cs.CLShow all

12 papers · 1 filter

cs.CL2021

Graph Algorithms for Multiparallel Word Alignment

Ayyoob Imani, Masoud Jalili Sabet, Lütfi Kerem Şenel +3

With the advent of end-to-end deep learning approaches in machine translation, interest in word alignments initially decreased; however, they have again become a focus of research…

cs.CL2021

Wine is Not v i n. -- On the Compatibility of Tokenizations Across Languages

Antonis Maronikolakis, Philipp Dufter, Hinrich Schütze

The size of the vocabulary is a central design choice in large pretrained language models, with respect to both performance and memory requirements. Typically, subword tokenization…

cs.CL2021

ParCourE: A Parallel Corpus Explorer for a Massively Multilingual Corpus

Ayyoob Imani, Masoud Jalili Sabet, Philipp Dufter +2

With more than 7000 languages worldwide, multilingual natural language processing (NLP) is essential both from an academic and commercial perspective. Researching typological prope…

cs.CL2021

Static Embeddings as Efficient Knowledge Bases?

Philipp Dufter, Nora Kassner, Hinrich Schütze

Recent research investigates factual knowledge stored in large pretrained language models (PLMs). Instead of structural knowledge base (KB) queries, masked sentences such as "Paris…

cs.CL2021

Position Information in Transformers: An Overview

Philipp Dufter, Martin Schmitt, Hinrich Schütze

Transformers are arguably the main workhorse in recent Natural Language Processing research. By definition a Transformer is invariant with respect to reordering of the input. Howev…

cs.CL2021

Multilingual LAMA: Investigating Knowledge in Multilingual Pretrained Language Models

Nora Kassner, Philipp Dufter, Hinrich Schütze

Recently, it has been found that monolingual English language models can be used as knowledge bases. Instead of structural knowledge base queries, masked sentences such as "Paris i…