1 paper
Jeongsoo Lee, Daeyong Kwon, Kyohoon Jin +3
Existing RAG benchmarks often overlook query difficulty, leading to inflated performance on simpler questions and unreliable evaluations. A robust benchmark dataset must satisfy th…