Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Rethinking Evaluation for LLM Hallucination Detection: A Desiderata, A New RAG-based Benchmark, New Insights
Wenbo Chen, Veena Padmanabhan, Tootiya Giyahchi +2
Hallucination, broadly referring to unfaithful, fabricated, or inconsistent content generated by LLMs, has wide-ranging implications. Therefore, a large body of effort has been dev…
cs.AI2026
Long-Horizon Plan Execution in Large Tool Spaces through Entropy-Guided Branching
Rongzhe Wei, Ge Shi, Min Cheng +5
Large Language Models (LLMs) have significantly advanced tool-augmented agents, enabling autonomous reasoning via API interactions. However, executing multi-step tasks within massi…
cs.AI2025
Uncertainty-aware Human Mobility Modeling and Anomaly Detection
Haomin Wen, Shurui Cao, Zeeshan Rasheed +2
Given the temporal GPS coordinates from a large set of human agents, how can we model their mobility behavior toward effective anomaly (e.g. bad-actor or malicious behavior) detect…