1 citations · 1 across the 3 of their papers we have counts for
4 papers
When Rubrics Fail: Hallucinations Reveal Blind Spots in Medical AI Evaluation
Griffin Farrow, Lily Sijia Li, Jack Johnson +4
Hallucinations can undermine clinician trust in LLMs, making it important that evaluation methods capture clinically relevant errors. Rubric-based evaluation has become the leading…
Learning Clinical-Trial Strategy: Offline Policy Training for Decision Agents
William Bolton, Philip Torr
Clinical development is sequential decision-making under uncertainty, where a sponsor must plan a portfolio of experiments from heterogeneous evidence. We study this setting by fra…
Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities
William Bolton, Philip Torr
Benchmarking the ability of AI scientists to generate novel ideas is notoriously difficult. Existing benchmarks in this field have made progress in evaluating scientific reasoning…
RAmBLA: A Framework for Evaluating the Reliability of LLMs as Assistants in the Biomedical Domain
William James Bolton, Rafael Poyiadzi, Edward R. Morrell +2
Large Language Models (LLMs) increasingly support applications in a wide range of domains, some with potential high societal impact such as biomedicine, yet their reliability in re…