1 paper
Qingchen Yu, Shichao Song, Ke Fang +5
As the application of Large Language Models (LLMs) expands, the demand for reliable evaluations increases. Existing LLM evaluation benchmarks primarily rely on static datasets, mak…