1 paper
Steve Han, Gilberto Titericz Junior, Tom Balough +1
This research introduces the Judge's Verdict Benchmark, a novel two-step methodology to evaluate Large Language Models (LLMs) as judges for response accuracy evaluation tasks. We a…