1 paper
Xinyu Hu, Mingqi Gao, Sen Hu +4
Some prior work has shown that LLMs perform well in NLG evaluation for different tasks. However, we discover that LLMs seem to confuse different evaluation criteria, which reduces…