1 paper
Aviya Maimon, Amir DN Cohen, Gal Vishne +2
Current evaluations of large language models (LLMs) rely heavily on a growing collection of benchmarks and on aggregate benchmark scores, yet it remains unclear what this compariso…