2 citations · 3 across the 6 of their papers we have counts for
1 paper · 1 filter
Yangsibo Huang, Milad Nasr, Anastasios Angelopoulos +10
It is now common to evaluate Large Language Models (LLMs) by having humans manually vote to evaluate model outputs, in contrast to typical benchmarks that evaluate knowledge or ski…