agentic search 1context management 1information retrieval 1long-horizon tasks 1web browsing agents 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.AI2026
FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents
Yuhao Zhang, O. Ozan Koyluoglu, Thejas Venkatesh +4
AI agents are increasingly deployed for professional investment research, yet no benchmark captures the complexity of the full investor workflow. Existing benchmarks mainly target…
cs.CL2026
Lost in the Maze: Overcoming Context Limitations in Long-Horizon Agentic Search
Howard Yen, Yoonsang Lee, Ashwin Paranjape +5
The paper introduces SLIM, a lightweight framework that separates search and browsing tools and periodically summarizes information to overcome context limits in long-horizon web‑a…
cs.CL2026
OpaqueToolsBench: Learning Nuances of Tool Behavior Through Interaction
Skyler Hallinan, Thejas Venkatesh, Xiang Ren +4
Tool-calling is essential for Large Language Model (LLM) agents to complete real-world tasks. While most existing benchmarks assume simple, perfectly documented tools, real-world t…