1 paper
Siyi Hao, Yidi Cao, Linhao Yu +2
With the wide application of large language models (LLMs) in real-world scenarios, the value implication of their outputs is crucial. However, existing evaluation benchmarks suffer…