1 paper · 1 filter
Junze Ye, Daniel Tawfik, Alex J. Goodell +3
Reference labels for machine-learning benchmarks are increasingly synthesized with LLM assistance, but their reliability remains underexamined. We audit MedCalc-Bench, a clinical b…