2 citations · 2 across the 1 of their papers we have counts for
3 papers
cs.AI2026★ 2 cited
Benchmark Test-Time Scaling of General LLM Agents
Xiaochuan Li, Ryan Ming, Pranav Setlur +6
LLM agents are increasingly expected to function as general-purpose systems capable of resolving open-ended user requests. While existing benchmarks focus on domain-aware environme…
cs.AI2026
Beneficial Reasoning Behaviors in Agentic Search and Effective Post-training to Obtain Them
Jiahe Jin, Abhijay Paladugu, Chenyan Xiong
Agentic search requires large language models (LLMs) to perform multi-step search to solve complex information-seeking tasks, imposing unique challenges on their reasoning capabili…
cs.IR2025
DeepResearchGym: A Free, Transparent, and Reproducible Evaluation Sandbox for Deep Research
João Coelho, Jingjie Ning, Jingyuan He +8
Deep research systems represent an emerging class of agentic information retrieval methods that generate comprehensive and well-supported reports to complex queries. However, most…