1 paper · 1 filter
Doan Nam Long Vu, Simone Balloccu
Trustworthy clinical AI requires that performance gains reflect genuine evidence integration rather than surface-level artifacts. We evaluate 12 open-weight vision-language models…