2 papers
cs.CL2026
RAGCap-Bench: Benchmarking Capabilities of LLMs in Agentic Retrieval Augmented Generation Systems
Jingru Lin, Chen Zhang, Stephen Y. Liu +1
Retrieval-Augmented Generation (RAG) mitigates key limitations of Large Language Models (LLMs)-such as factual errors, outdated knowledge, and hallucinations-by dynamically retriev…
cs.AI2026
SourceBench: Can AI Answers Reference Quality Web Sources?
Hexi Jin, Stephen Liu, Yuheng Li +2
Large language models (LLMs) increasingly answer queries by citing web sources, but existing evaluations emphasize answer correctness rather than evidence quality. We introduce Sou…