3 papers
cs.CL2026
Soft-Prompt Tuning for Fair and Efficient LLM Benchmark Evaluation
Selen Erkan, Bastian Boll, Kristian Kersting +2
Benchmark scores often misrepresent a large language model's (LLM's) knowledge, because they rely, e.g., on the model's ability to follow specific formatting requirements. This esp…
cs.CV2025
False Promises in Medical Imaging AI? Assessing Validity of Outperformance Claims
Evangelia Christodoulou, Annika Reinke, Pascaline Andrè +23
Performance comparisons are fundamental in medical imaging Artificial Intelligence (AI) research, often driving claims of superiority based on relative improvements in common perfo…
cs.CV2025
SURE-VQA: Systematic Understanding of Robustness Evaluation in Medical VQA Tasks
Kim-Celine Kahl, Selen Erkan, Jeremias Traub +4
Vision-Language Models (VLMs) have great potential in medical tasks, like Visual Question Answering (VQA), where they could act as interactive assistants for both patients and clin…