1 paper · 1 filter
William Philipp, Finn Fassbender, Thorsten Langer +16
Open-response evaluation provides stronger clinical validity than multiple-choice benchmarks but creates a scoring bottleneck that motivates automated LLM-asa-Judge approaches. Whe…