12 papers · 1 filter
Graph Algorithms for Multiparallel Word Alignment
Ayyoob Imani, Masoud Jalili Sabet, Lütfi Kerem Şenel +3
With the advent of end-to-end deep learning approaches in machine translation, interest in word alignments initially decreased; however, they have again become a focus of research…
Wine is Not v i n. -- On the Compatibility of Tokenizations Across Languages
Antonis Maronikolakis, Philipp Dufter, Hinrich Schütze
The size of the vocabulary is a central design choice in large pretrained language models, with respect to both performance and memory requirements. Typically, subword tokenization…
ParCourE: A Parallel Corpus Explorer for a Massively Multilingual Corpus
Ayyoob Imani, Masoud Jalili Sabet, Philipp Dufter +2
With more than 7000 languages worldwide, multilingual natural language processing (NLP) is essential both from an academic and commercial perspective. Researching typological prope…
Static Embeddings as Efficient Knowledge Bases?
Philipp Dufter, Nora Kassner, Hinrich Schütze
Recent research investigates factual knowledge stored in large pretrained language models (PLMs). Instead of structural knowledge base (KB) queries, masked sentences such as "Paris…
Position Information in Transformers: An Overview
Philipp Dufter, Martin Schmitt, Hinrich Schütze
Transformers are arguably the main workhorse in recent Natural Language Processing research. By definition a Transformer is invariant with respect to reordering of the input. Howev…
Multilingual LAMA: Investigating Knowledge in Multilingual Pretrained Language Models
Nora Kassner, Philipp Dufter, Hinrich Schütze
Recently, it has been found that monolingual English language models can be used as knowledge bases. Instead of structural knowledge base queries, masked sentences such as "Paris i…