activity
20242026
collaborators

10 papers

cs.AI2026

Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence

Jakub Pokrywka, Łukasz Grzybowski, Antoni Lasik +3

We introduce a Polish-language medical visual question answering (VQA) benchmark, built from Polish Board Certification Examination questions for licensed physicians and dentists p…

cs.CL2026

Jako Tako or Fluent? Presenting PoVisLE: A Polish Vision-Language Evaluation

Anna Kołos, Grzegorz Statkiewicz, Karolina Seweryn +3

Vision-language models (VLMs) have achieved strong performance on tasks such as image captioning, visual question answering, and image-to-text generation. However, they are predomi…

cs.CL2026

Reassessing High-Performing LLMs on Polish Medical Exams: True Competence or Bias-Driven Performance?

Antoni Lasik, Jakub Pokrywka, Łukasz Grzybowski +7

Large language models (LLMs) in medicine are mainly evaluated using multiple-choice question answering (MCQA), which can overestimate real clinical ability due to guessing strategi…

cs.CL2026

Multilingual Refusal Alignment for Safer Large Language Models

Aleksandra Krasnodębska, Wojciech Kusa, Aldo Lipani

As Large Language Models (LLMs) are deployed globally, ensuring their safety and alignment across multiple languages becomes paramount. However, safety behaviors often vary unpredi…

cs.IR2026

The LLM Effect on IR Benchmarks: A Meta-Analysis of Effectiveness, Baselines, and Contamination

Moritz Staudinger, Wojciech Kusa, Allan Hanbury

Benchmark collections have long enabled controlled comparison and cumulative progress in Information Retrieval (IR). However, prior meta-analyses have shown that reported effective…

cs.CL2026

Annotation-Efficient Vision-Language Model Adaptation to the Polish Language Using the LLaVA Framework

Grzegorz Statkiewicz, Alicja Dobrzeniecka, Karolina Seweryn +5

Most vision-language models (VLMs) are trained on English-centric data, limiting their performance in other languages and cultural contexts. This restricts their usability for non-…