1 citations · 1 across the 8 of their papers we have counts for
1 paper · 1 filter
Yuhang Zhou, Xutian Chen, Yixin Cao +8
Recent progress in large language models (LLMs) has outpaced the development of effective evaluation methods. Traditional benchmarks rely on task-specific metrics and static datase…