2 papers
cs.AI2026
Long-Horizon State Tracking in LLMs: Executing MD5 through a Deep Sequence of Dependent Tool Calls
Dheeraj Mohandas Pai, Lu Xian
Long-horizon tasks remain uncommon in large language model (LLM) evaluation, and for a reason: when each step depends on the last, per-step accuracy that looks excellent in isolati…
cs.AI2026
FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud
Dheeraj Mohandas Pai, Lu Xian
Conversational agents now act for end users through tools while holding access to customer databases and internal policy documents that a caller can reach through dialogue alone. B…