1 paper · 1 filter
Harshada Badave, Santosh Borse, Andrea Gomez +6
Large Language Models (LLMs) are increasingly deployed as autonomous agents that reason, use tools, and act over multiple steps. Yet most hallucination benchmarks still evaluate on…