3 papers
cs.CL2026
ForMaT: Dataset for Visually-Grounded Multilingual PDF Translation
Michał Ciesiółka, Dawid Wiśniewski, Adrian Charkiewicz +1
We present ForMaT (Format-Preserving Multilingual Translation), a parallel corpus of 3,956 PDFs across 15 language pairs that preserves original layout metadata proposed for multim…
cs.CL2025
LLMzSzŁ: a comprehensive LLM benchmark for Polish
Krzysztof Jassem, Michał Ciesiółka, Filip Graliński +5
This article introduces the first comprehensive benchmark for the Polish language at this scale: LLMzSzŁ (LLMs Behind the School Desk). It is based on a coherent collection of Poli…
cs.CL2024
Polish-English medical knowledge transfer: A new benchmark and results
Łukasz Grzybowski, Jakub Pokrywka, Michał Ciesiółka +2
Large Language Models (LLMs) have demonstrated significant potential in handling specialized tasks, including medical problem-solving. However, most studies predominantly focus on…