50 citations · 145 across the 47 of their papers we have counts for
Showing 2026 · cs.AIShow all
2 papers · 2 filters
cs.AI2026
Rethinking Evaluation for LLM Hallucination Detection: A Desiderata, A New RAG-based Benchmark, New Insights
Wenbo Chen, Veena Padmanabhan, Tootiya Giyahchi +2
Hallucination, broadly referring to unfaithful, fabricated, or inconsistent content generated by LLMs, has wide-ranging implications. Therefore, a large body of effort has been dev…
cs.AI2026
Long-Horizon Plan Execution in Large Tool Spaces through Entropy-Guided Branching
Rongzhe Wei, Ge Shi, Min Cheng +5
Large Language Models (LLMs) have significantly advanced tool-augmented agents, enabling autonomous reasoning via API interactions. However, executing multi-step tasks within massi…