19 papers
RetailBench: Evaluating Long-Horizon Autonomous Decision-Making and Strategy Stability of LLM Agents in Realistic Retail Environments
Linghua Zhang, Jun Wang, Jingtong Wu +1
Large language model (LLM) agents have made rapid progress on short-horizon, well-scoped tasks, yet their ability to sustain coherent decisions in dynamic long-horizon environments…
RetailBench: Benchmarking long horizon reasoning and coherent decision making of LLM agents in realistic retail environments
Linghua Zhang, Jun Wang, Jingtong Wu +1
Large language model (LLM) agents have made rapid progress on short-horizon, well-scoped tasks, yet their ability to sustain coherent decisions in dynamic long-horizon environments…
Communication Policy Evolution for Proactive LLM Agents
Xinbei Ma, Jiyang Qiu, Yao Yao +10
LLM agents have rapidly evolved into autonomous systems, yet a persistent information gap remains between users and agents: communication is costly, while users' identical preferen…
Retrospective Progress-Aware Self-Refinement for LLM Agent Training
Xinbei Ma, Congmin Zheng, Jiyang Qiu +10
LLM-based agents trained with reinforcement learning optimize step-wise action prediction but lack metacognitive awareness of task progress, inducing a gap that hinders long-horizo…
Holder Policy Optimisation
Yuxiang Chen, Dingli Liang, Yihang Chen +8
Group Relative Policy Optimisation (GRPO) enhances large language models by estimating advantages across a group of sampled trajectories. However, mapping these trajectory-level ad…
Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents
Jingxing Wang, Chenyu Zhou, Zhihui Fu +4
Additional test-time compute can give LLM agents access to more past experience, yet expanding the context or adding rollouts does not necessarily yield greater agent capability. W…