2 citations · 2 across the 2 of their papers we have counts for
3 papers
Uncertainty Is Not a Safety Net for Clinical VQA, but Can It Anticipate Model Failure?
Arnisa Fazla, Alberto Testoni, Ameen Abu-Hanna +2
Safe deployment of clinical vision-language models (VLMs) requires reliable uncertainty estimation (UE): a signal indicating when predictions should be trusted or escalated to a cl…
Robustness of Misinformation Classification Systems to Adversarial Examples Through BeamAttack
Arnisa Fazla, Lucas Krauter, David Guzman Piedrahita +1
We extend BeamAttack, an adversarial attack algorithm designed to evaluate the robustness of text classification systems through word-level modifications guided by beam search. Our…
Machine Translation Meta Evaluation through Translation Accuracy Challenge Sets
Nikita Moghe, Arnisa Fazla, Chantal Amrhein +5
Recent machine translation (MT) metrics calibrate their effectiveness by correlating with human judgement but without any insights about their behaviour across different error type…