4 papers
Can Continual Pre-training Bridge the Performance Gap between General-purpose and Specialized Language Models in the Medical Domain?
Niclas Doll, Jasper Schulze Buschhoff, Shalaka Satheesh +3
This paper narrows the performance gap between small, specialized models and significantly larger general-purpose models through domain adaptation via continual pre-training and me…
Textual Data Bias Detection and Mitigation -- An Extensible Pipeline with Experimental Evaluation
Rebekka Görge, Sujan Sai Gannamaneni, Tabea Naeven +10
Textual data used to train large language models (LLMs) exhibits multifaceted bias manifestations encompassing harmful language and skewed demographic distributions. Regulations su…
Teuken-7B-Base & Teuken-7B-Instruct: Towards European LLMs
Mehdi Ali, Michael Fromm, Klaudia Thellmann +38
We present two multilingual LLMs, Teuken 7B-base and Teuken 7B-instruct, designed to embrace Europe's linguistic diversity by supporting all 24 official languages of the European U…
Data Processing for the OpenGPT-X Model Family
Nicolo' Brandizzi, Hammam Abdelwahab, Anirban Bhowmick +19
This paper presents a comprehensive overview of the data preparation pipeline developed for the OpenGPT-X project, a large-scale initiative aimed at creating open and high-performa…