8 citations · 9 across the 5 of their papers we have counts for
4 papers · 1 filter
The German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models
Lukas Gienapp, Christopher Schröder, Stefan Schweter +5
Large language model development relies on large-scale training corpora, yet most contain data of unclear licensing status, limiting the development of truly open models. This prob…
Entities, Dates, and Languages: Zero-Shot on Historical Texts with T0
Francesco De Toni, Christopher Akiki, Javier de la Rosa +4
In this work, we explore whether the recently demonstrated zero-shot abilities of the T0 model extend to Named Entity Recognition for out-of-distribution languages and time periods…
Tracking Discourse Influence in Darknet Forums
Christopher Akiki, Lukas Gienapp, Martin Potthast
This technical report documents our efforts in addressing the tasks set forth by the 2021 AMoC (Advanced Modelling of Cyber Criminal Careers) Hackathon. Our main contribution is a…
BERTian Poetics: Constrained Composition with Masked LMs
Christopher Akiki, Martin Potthast
Masked language models have recently been interpreted as energy-based sequence models that can be generated from using a Metropolis--Hastings sampler. This short paper demonstrates…