collaborators

6 papers

cs.AI2026

Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence

Jakub Pokrywka, Łukasz Grzybowski, Antoni Lasik +3

We introduce a Polish-language medical visual question answering (VQA) benchmark, built from Polish Board Certification Examination questions for licensed physicians and dentists p…

cs.SD2026

A Benchmark for Audio Reasoning Capabilities of Multimodal Large Language Models

Iwona Christop, Mateusz Czyżnikiewicz, Paweł Skórzewski +4

The present benchmarks for testing the audio modality of multimodal large language models concentrate on testing various audio tasks such as speaker diarization or gender identific…

cs.CL2025

Preservation of Language Understanding Capabilities in Speech-aware Large Language Models

Marek Kubis, Paweł Skórzewski, Iwona Christop +4

The paper presents C3T (Cross-modal Capabilities Conservation Test), a new benchmark for assessing the performance of speech-aware large language models. The benchmark utilizes tex…

cs.CL2025

ClonEval: An Open Voice Cloning Benchmark

Iwona Christop, Tomasz Kuczyński, Marek Kubis

We present a novel benchmark for voice cloning text-to-speech models. The benchmark consists of an evaluation protocol, an open-source library for assessing the performance of voic…

cs.CL2025

Polish-English medical knowledge transfer: A new benchmark and results

Łukasz Grzybowski, Jakub Pokrywka, Michał Ciesiółka +2

Large Language Models (LLMs) have demonstrated significant potential in handling specialized tasks, including medical problem-solving. However, most studies predominantly focus on…

cs.CL2025

LLMzSzŁ: a comprehensive LLM benchmark for Polish

Krzysztof Jassem, Michał Ciesiółka, Filip Graliński +5

This article introduces the first comprehensive benchmark for the Polish language at this scale: LLMzSzŁ (LLMs Behind the School Desk). It is based on a coherent collection of Pol…