1 paper · 1 filter
Rubing Chen, Jiaxin Wu, Jian Wang +5
The increasing demand for domain-specific evaluation of large language models (LLMs) has led to the development of numerous benchmarks. These efforts often adhere to the principle…