6 papers
Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence
Jakub Pokrywka, Łukasz Grzybowski, Antoni Lasik +3
We introduce a Polish-language medical visual question answering (VQA) benchmark, built from Polish Board Certification Examination questions for licensed physicians and dentists p…
A Benchmark for Audio Reasoning Capabilities of Multimodal Large Language Models
Iwona Christop, Mateusz Czyżnikiewicz, PaweŠSkórzewski +4
The present benchmarks for testing the audio modality of multimodal large language models concentrate on testing various audio tasks such as speaker diarization or gender identific…
Preservation of Language Understanding Capabilities in Speech-aware Large Language Models
Marek Kubis, PaweŠSkórzewski, Iwona Christop +4
The paper presents C3T (Cross-modal Capabilities Conservation Test), a new benchmark for assessing the performance of speech-aware large language models. The benchmark utilizes tex…
ClonEval: An Open Voice Cloning Benchmark
Iwona Christop, Tomasz KuczyÅski, Marek Kubis
We present a novel benchmark for voice cloning text-to-speech models. The benchmark consists of an evaluation protocol, an open-source library for assessing the performance of voic…
Polish-English medical knowledge transfer: A new benchmark and results
Åukasz Grzybowski, Jakub Pokrywka, MichaÅ CiesióÅka +2
Large Language Models (LLMs) have demonstrated significant potential in handling specialized tasks, including medical problem-solving. However, most studies predominantly focus on…
LLMzSzÅ: a comprehensive LLM benchmark for Polish
Krzysztof Jassem, MichaÅ CiesióÅka, Filip GraliÅski +5
This article introduces the first comprehensive benchmark for the Polish language at this scale: LLMzSzÅ (LLMs Behind the School Desk). It is based on a coherent collection of Pol…