1 citations · 1 across the 1 of their papers we have counts for
1 paper
Steve Han, Gilberto Titericz Junior, Tom Balough +1
This research introduces the Judge's Verdict Benchmark, a novel two-step methodology to evaluate Large Language Models (LLMs) as judges for response accuracy evaluation tasks. We a…