6 papers · 1 filter
ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents
Shuhan Xue, Jianyuan Zhong, Ziyuan Nan +10
We introduce and release ScienceBuddy, an interactive scientific research workspace that brings continually improving scientific agents into researchers' everyday workflows. Scienc…
RAP: Research Attention Prediction Reveals Target-Conditioned Evidence Acquisition Biases
Yingqian Wu, Jingcong Liang, Siyuan Wang +4
Large language models (LLMs) increasingly act as research agents, yet their ability to track shifts in research attention is difficult to evaluate because reviews and research idea…
Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses
Zhaochen Yu, Yingcheng Wu, Zhenfei Yin +5
Recursive self-improvement (RSI) remains hard in long-horizon tasks, where growing histories obscure the task state and misalign skill invocation. We introduce Recuris, a recursive…
Strategic Exploitation in LLM Agent Markets: A Simulation Framework for E-Commerce Trust
Shijun Lei, Quang Nguyen, Swapneel S Mehta +7
Agent-based modeling (ABM) has long been used in economics to study human behavior, and large language model (LLM) agents now enable new forms of social and economic simulation. Wh…
Rethinking the Role of Entropy in Optimizing Tool-Use Behaviors for Large Language Model Agents
Zeping Li, Hongru Wang, Yiwen Zhao +7
Tool-using agents based on Large Language Models (LLMs) excel in tasks such as mathematical reasoning and multi-hop question answering. However, in long trajectories, agents often…
Why Reasoning Fails to Plan: A Planning-Centric Analysis of Long-Horizon Decision Making in LLM Agents
Zehong Wang, Fang Wu, Hongru Wang +8
Large language model (LLM)-based agents exhibit strong step-by-step reasoning capabilities over short horizons, yet often fail to sustain coherent behavior over long planning horiz…