2 papers
cs.AI2026
Scalable Stewardship of an LLM-Assisted Clinical Benchmark with Physician Oversight
Junze Ye, Daniel Tawfik, Alex J. Goodell +3
Reference labels for machine-learning benchmarks are increasingly synthesized with LLM assistance, but their reliability remains underexamined. We audit MedCalc-Bench, a clinical b…
cs.AI2025
Scaling Clinician-Grade Feature Generation from Clinical Notes with Multi-Agent Language Models
Jiayi Wang, Jacqueline Jil Vallon, Nikhil V. Kotha +8
Developing accurate clinical prediction models is often bottlenecked by the difficulty of deriving meaningful structured features from unstructured EHR notes, a process that tradit…