collaborators

5 papers

cs.CL2026

SciLaD: A Large-Scale, Transparent, Reproducible Dataset for Natural Scientific Language Processing

Luca Foppiano, Sotaro Takeshita, Pedro Ortiz Suarez +6

SciLaD is a novel, large-scale dataset of scientific language constructed entirely using open-source frameworks and publicly available data sources. It comprises a curated English…

cs.IR2026

Colour Contrast on the Web: A WCAG 2.1 Level AA Compliance Audit of Common Crawl's Top 500 Domains

Thom Vaughan, Pedro Ortiz Suarez

We present a large-scale automated audit of WCAG 2.1/2.2 Level AA colour contrast compliance across the 500 most frequently crawled registered domains in Common Crawl's CC-MAIN-202…

cs.CL2025

Teuken-7B-Base & Teuken-7B-Instruct: Towards European LLMs

Mehdi Ali, Michael Fromm, Klaudia Thellmann +38

We present two multilingual LLMs, Teuken 7B-base and Teuken 7B-instruct, designed to embrace Europe's linguistic diversity by supporting all 24 official languages of the European U…

cs.CL2025

Data Processing for the OpenGPT-X Model Family

Nicolo' Brandizzi, Hammam Abdelwahab, Anirban Bhowmick +19

This paper presents a comprehensive overview of the data preparation pipeline developed for the OpenGPT-X project, a large-scale initiative aimed at creating open and high-performa…

cs.CL2025

Table Understanding and (Multimodal) LLMs: A Cross-Domain Case Study on Scientific vs. Non-Scientific Data

Ekaterina Borisova, Fabio Barth, Nils Feldhus +5

Tables are among the most widely used tools for representing structured data in research, business, medicine, and education. Although LLMs demonstrate strong performance in downstr…