1 paper · 1 filter
Jeongsoo Lee, Daeyong Kwon, Kyohoon Jin +3
Existing RAG benchmarks often overlook query difficulty, leading to inflated performance on simpler questions and unreliable evaluations. A robust benchmark dataset must satisfy th…