2 citations · 2 across the 3 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026★ 2 cited
Benchmark Test-Time Scaling of General LLM Agents
Xiaochuan Li, Ryan Ming, Pranav Setlur +6
LLM agents are increasingly expected to function as general-purpose systems capable of resolving open-ended user requests. While existing benchmarks focus on domain-aware environme…
cs.AI2025
ResearchArena: Benchmarking Large Language Models' Ability to Collect and Organize Information as Research Agents
Hao Kang, Chenyan Xiong
Large language models (LLMs) excel across many natural language processing tasks but face challenges in domain-specific, analytical tasks such as conducting research surveys. This…