2 papers
cs.CL2026
LoHoSearch: Benchmarking Long-Horizon Search Agents Beyond the Human Difficulty Ceiling
Jiarui Zhao, Rongzhi Zhang, Lingchuan Liu +3
Search agent benchmarks exemplified by BrowseComp have rapidly saturated over the past year, with the strongest models surpassing 90% accuracy. Since these benchmarks are predomina…
cs.IR2024
A Multi-Source Retrieval Question Answering Framework Based on RAG
Ridong Wu, Shuhong Chen, Xiangbiao Su +3
With the rapid development of large-scale language models, Retrieval-Augmented Generation (RAG) has been widely adopted. However, existing RAG paradigms are inevitably influenced b…