Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Evaluation of retrieval-based QA on QUEST-LOFT
Nathan Scales, Nathanael Schärli, Olivier Bousquet
Despite the popularity of retrieval-augmented generation (RAG) as a solution for grounded QA in both academia and industry, current RAG methods struggle with questions where the ne…
cs.CL2024
When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards
Norah Alzahrani, Hisham Abdullah Alyahya, Yazeed Alnumay +9
Large Language Model (LLM) leaderboards based on benchmark rankings are regularly used to guide practitioners in model selection. Often, the published leaderboard rankings are take…