activity
20242026
collaborators

5 papers

cs.LG2026

Modalities, a PyTorch-native Framework For Large-scale LLM Training and Research

Max Lübbering, Timm Ruland, Richard Rutmann +8

Today's LLM (pre-) training and research workflows typically allocate a significant amount of compute to large-scale ablation studies. Despite the substantial compute costs of thes…

cs.CL2025

Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models

Mehdi Ali, Manuel Brack, Max Lübbering +15

High-quality multilingual training data is essential for effectively pretraining large language models (LLMs). Yet, the availability of suitable open-source multilingual datasets r…

cs.CL2024

Towards Multilingual LLM Evaluation for European Languages

Klaudia Thellmann, Bernhard Stadler, Michael Fromm +8

The rise of Large Language Models (LLMs) has revolutionized natural language processing across numerous languages and tasks. However, evaluating LLM performance in a consistent and…

cs.CL2024

Data Processing for the OpenGPT-X Model Family

Nicolo' Brandizzi, Hammam Abdelwahab, Anirban Bhowmick +19

This paper presents a comprehensive overview of the data preparation pipeline developed for the OpenGPT-X project, a large-scale initiative aimed at creating open and high-performa…

cs.CL2024

Teuken-7B-Base & Teuken-7B-Instruct: Towards European LLMs

Mehdi Ali, Michael Fromm, Klaudia Thellmann +38

We present two multilingual LLMs, Teuken 7B-base and Teuken 7B-instruct, designed to embrace Europe's linguistic diversity by supporting all 24 official languages of the European U…