Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
LexAgentHallu: A Hierarchical Benchmark for Profiling Hallucinations in Legal Agents
Yujin Zhou, Mingxuan Zheng, Chuxue Cao +4
As large language models are increasingly deployed as tool-augmented legal agents, they introduce agentic hallucinations where tool-call and reasoning errors cascade into fabricate…
cs.AI2026
SearchAuditor: Auditing and Attributing Failures in Long-Horizon Search Agents
Zhixiang Liang, Yifei Liu, Yidan Huang +5
Deep search agents tackle challenging questions through long-horizon web interactions, a process that is both complex and fragile: small reasoning errors may propagate through long…