4 citations · 11 across the 11 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Evaluating Metrics for Safety with LLM-as-Judges
Kester Clegg, Richard Hawkins, Ibrahim Habli +1
LLMs (Large Language Models) are increasingly used in text processing pipelines to intelligently respond to a variety of inputs and generation tasks. This raises the possibility of…
cs.CL2025
WER is Unaware: Assessing How ASR Errors Distort Clinical Understanding in Patient Facing Dialogue
Zachary Ellis, Jared Joselowitz, Yash Deo +7
As Automatic Speech Recognition (ASR) is increasingly deployed in clinical dialogue, standard evaluations still rely heavily on Word Error Rate (WER). This paper challenges that st…