1 paper
Shixin Fang, Jiachen Wo, Wenjuan Qin +2
Large language model (LLM) evaluation spans diverse tasks and benchmarks, yet evidence remains organized around tasks rather than the capabilities they probe. This fragmentation li…