2 papers
cs.AI2026
GuardianAgentBench: Where Agents Fail and How to Guard Them
Vishal Ishwar Naik, Chenyu Xu, Donna Dong +5
As large language model agents increasingly operate autonomously with access to tools and external environments, ensuring their safe and reliable behavior becomes critical. We pres…
cs.CL2025
Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards
Manveer Singh Tamber, Forrest Sheng Bao, Chenyu Xu +7
Retrieval-augmented generation (RAG) aims to reduce hallucinations by grounding responses in external context, yet large language models (LLMs) still frequently introduce unsupport…