Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Yubo Wang, Xueguang Ma, Ge Zhang +14
In the age of large-scale language models, benchmarks like the Massive Multitask Language Understanding (MMLU) have been pivotal in pushing the boundaries of what AI can achieve in…
cs.CL2024
ReFeR: Improving Evaluation and Reasoning through Hierarchy of Models
Yaswanth Narsupalli, Abhranil Chandra, Sreevatsa Muppirala +2
Assessing the quality of outputs generated by generative models, such as large language models and vision language models, presents notable challenges. Traditional methods for eval…