15 citations · 16 across the 3 of their papers we have counts for
3 papers
cs.CL2024
Symmetric Dot-Product Attention for Efficient Training of BERT Language Models
Martin Courtois, Malte Ostendorff, Leonhard Hennig +1
Initially introduced as a machine translation model, the Transformer architecture has now become the foundation for modern deep learning architecture, with applications in a wide r…
cs.DL2024★ 1 cited
Toward FAIR Semantic Publishing of Research Dataset Metadata in the Open Research Knowledge Graph
Raia Abu Ahmad, Jennifer D'Souza, Matthäus Zloch +5
Search engines these days can serve datasets as search results. Datasets get picked up by search technologies based on structured descriptions on their official web pages, informed…
cs.CL2023★ 15 cited
Efficient Language Model Training through Cross-Lingual and Progressive Transfer Learning
Malte Ostendorff, Georg Rehm
Most Transformer language models are primarily pretrained on English text, limiting their use for other languages. As the model sizes grow, the performance gap between English and…