10 papers
Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence
Jakub Pokrywka, Łukasz Grzybowski, Antoni Lasik +3
We introduce a Polish-language medical visual question answering (VQA) benchmark, built from Polish Board Certification Examination questions for licensed physicians and dentists p…
Jako Tako or Fluent? Presenting PoVisLE: A Polish Vision-Language Evaluation
Anna Kołos, Grzegorz Statkiewicz, Karolina Seweryn +3
Vision-language models (VLMs) have achieved strong performance on tasks such as image captioning, visual question answering, and image-to-text generation. However, they are predomi…
Reassessing High-Performing LLMs on Polish Medical Exams: True Competence or Bias-Driven Performance?
Antoni Lasik, Jakub Pokrywka, Åukasz Grzybowski +7
Large language models (LLMs) in medicine are mainly evaluated using multiple-choice question answering (MCQA), which can overestimate real clinical ability due to guessing strategi…
Multilingual Refusal Alignment for Safer Large Language Models
Aleksandra KrasnodÄbska, Wojciech Kusa, Aldo Lipani
As Large Language Models (LLMs) are deployed globally, ensuring their safety and alignment across multiple languages becomes paramount. However, safety behaviors often vary unpredi…
The LLM Effect on IR Benchmarks: A Meta-Analysis of Effectiveness, Baselines, and Contamination
Moritz Staudinger, Wojciech Kusa, Allan Hanbury
Benchmark collections have long enabled controlled comparison and cumulative progress in Information Retrieval (IR). However, prior meta-analyses have shown that reported effective…
Annotation-Efficient Vision-Language Model Adaptation to the Polish Language Using the LLaVA Framework
Grzegorz Statkiewicz, Alicja Dobrzeniecka, Karolina Seweryn +5
Most vision-language models (VLMs) are trained on English-centric data, limiting their performance in other languages and cultural contexts. This restricts their usability for non-…