1 paper
William Philipp, Finn Fassbender, Thorsten Langer +16
Open-response evaluation provides stronger clinical validity than multiple-choice benchmarks but creates a scoring bottleneck that motivates automated LLM-asa-Judge approaches. Whe…