1 paper
Juntao Wu, Wei Wen, Xianting Huang +4
Evaluating the exhaustive search capabilities of large language models (LLMs) is plagued by a fundamental paradox: verifying completeness requires complete ground truth, yet high-e…