2 papers
cs.AI2026
Scalable Stewardship of an LLM-Assisted Clinical Benchmark with Physician Oversight
Junze Ye, Daniel Tawfik, Alex J. Goodell +3
Reference labels for machine-learning benchmarks are increasingly synthesized with LLM assistance, but their reliability remains underexamined. We audit MedCalc-Bench, a clinical b…
cs.AI2025
VeriFact: Verifying Facts in LLM-Generated Clinical Text with Electronic Health Records
Philip Chung, Akshay Swaminathan, Alex J. Goodell +26
Methods to ensure factual accuracy of text generated by large language models (LLM) in clinical medicine are lacking. VeriFact is an artificial intelligence system that combines re…