13 citations · 13 across the 1 of their papers we have counts for
1 paper · 1 filter
Jiaju Lin, Haoran Zhao, Aochi Zhang +3
With ChatGPT-like large language models (LLM) prevailing in the community, how to evaluate the ability of LLMs is an open question. Existing evaluation methods suffer from followin…