1 paper
Jin Liu, Qingquan Li, Wenlong Du
In current benchmarks for evaluating large language models (LLMs), there are issues such as evaluation content restriction, untimely updates, and lack of optimization guidance. In…