6 papers · 1 filter
A Benchmark for Deep Information Synthesis
Debjit Paul, Daniel Murphy, Milan Gritta +14
Large language model (LLM)-based agents are increasingly used to solve complex tasks involving tool use, such as web browsing, code execution, and data analysis. However, current e…
Proof-of-Use: Mitigating Tool-Call Hacking in Deep Research Agents
SHengjie Ma, Chenlong Deng, Jiaxin Mao +5
While reinforcement learning (RL) enhances their ability to plan and reason across retrieval steps, we identify a critical failure mode in this setting: Tool-Call Hacking. Unlike e…
SAGE: Strategy-Adaptive Generation Engine for Query Rewriting
Teng Wang, Hailei Gong, Changwang Zhang +1
Query rewriting is pivotal for enhancing dense retrieval, yet current methods demand large-scale supervised data or suffer from inefficient reinforcement learning (RL) exploration.…
Efficient Agents: Building Effective Agents While Reducing Cost
Ningning Wang, Xavier Hu, Pai Liu +11
The remarkable capabilities of Large Language Model (LLM)-driven agents have enabled sophisticated systems to tackle complex, multi-step tasks, but their escalating costs threaten…
OAgents: An Empirical Study of Building Effective Agents
He Zhu, Tianrui Qin, King Zhu +21
Recently, Agentic AI has become an increasingly popular research field. However, we argue that current agent research practices lack standardization and scientific rigor, making it…
Scaling Test-time Compute for LLM Agents
King Zhu, Hanhao Li, Siwei Wu +12
Scaling test time compute has shown remarkable success in improving the reasoning abilities of large language models (LLMs). In this work, we conduct the first systematic explorati…