1 paper · 1 filter
Yi Xu, Laura Ruis, Tim Rocktäschel +1
Automatic evaluation methods based on large language models (LLMs) are emerging as the standard tool for assessing the instruction-following abilities of LLM-based agents. The most…