1 paper
Sahel Sharifymoghaddam, Yijun Ge, Jimmy Lin
Retrieval benchmarks for large language models (LLMs) should reflect the long, reasoning-intensive queries typical of retrieval-augmented generation (RAG). We present a systematic…