1 paper
Sebastian Heineking, Jonas Probst, Daniel Steinbach +2
Evaluating the output of generative large language models (LLMs) is challenging and difficult to scale. Many evaluations of LLMs focus on tasks such as single-choice question-answe…