1 paper
Kehua Feng, Keyan Ding, Hongzhi Tan +8
Reliable evaluation of large language models (LLMs) is impeded by two key challenges: objective metrics often fail to reflect human perception of natural language, and exhaustive h…