2 papers
cs.AI2026
RetailBench: Evaluating Long-Horizon Autonomous Decision-Making and Strategy Stability of LLM Agents in Realistic Retail Environments
Linghua Zhang, Jun Wang, Jingtong Wu +1
Large language model (LLM) agents have made rapid progress on short-horizon, well-scoped tasks, yet their ability to sustain coherent decisions in dynamic long-horizon environments…
cs.AI2026
RetailBench: Benchmarking long horizon reasoning and coherent decision making of LLM agents in realistic retail environments
Linghua Zhang, Jun Wang, Jingtong Wu +1
Large language model (LLM) agents have made rapid progress on short-horizon, well-scoped tasks, yet their ability to sustain coherent decisions in dynamic long-horizon environments…