3 papers
cs.CR2026
Trivial Prompt Reframing Bypasses Safety Guardrails in GoogleÅ MedGemma-4B
Avi-ad Avraam Buskila
Open-weight medical language models are increasingly used as the base of patient-facing and clinician-support applications. Their model cards prohibit specific behaviors -- recomme…
cs.CL2026
Domain Fine-Tuning vs. Retrieval-Augmented Generation for Medical Multiple-Choice Question Answering: A Controlled Comparison at the 4B-Parameter Scale
Avi-ad Avraam Buskila
Practitioners deploying small open-weight large language models (LLMs) for medical question answering face a recurring design choice: invest in a domain-fine-tuned model, or keep a…
cs.IR2026
Evaluating Small Open LLMs for Medical Question Answering: A Practical Framework
Avi-ad Avraam Buskila
Incorporating large language models (LLMs) in medical question answering demands more than high average accuracy: a model that returns substantively different answers each time it…