1 paper
Bing Zhang, Mikio Takeuchi, Ryo Kawahara +5
The advancement of large language models (LLMs) has led to a greater challenge of having a rigorous and systematic evaluation of complex tasks performed, especially in enterprise a…