3 papers
cs.CL2025
MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge
Jie He, Nan Hu, Wanqiu Long +2
Large language models (LLMs) have demonstrated impressive capabilities in various reasoning tasks but face significant challenges with complex, knowledge-intensive multi-hop querie…
cs.CR2025
RAG Safety: Exploring Knowledge Poisoning Attacks to Retrieval-Augmented Generation
Tianzhe Zhao, Jiaoyan Chen, Yanchi Ru +4
Retrieval-Augmented Generation (RAG) enhances large language models (LLMs) by retrieving external data to mitigate hallucinations and outdated knowledge issues. Benefiting from the…
cs.CL2025
Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges
Hongru Wang, Wenyu Huang, Yufei Wang +7
Existing benchmarks that assess Language Models (LMs) as Language Agents (LAs) for tool use primarily focus on stateless, single-turn interactions or partial evaluations, such as t…