5 papers
PACEvolve++: Improving Test-time Learning for Evolutionary Search Agents
Minghao Yan, Bo Peng, Benjamin Coleman +11
Large language models have become drivers of evolutionary search, but most systems rely on a fixed, prompt-elicited policy to sample next candidates. This limits adaptation in prac…
PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Research
Tingjia Miao, Wenkai Jin, Muhua Zhang +19
The paradigm of agentic science requires AI systems to conduct robust reasoning and engage in long-horizon, autonomous exploration. However, current scientific benchmarks remain co…
Web Agents Should Use Typed Actions Instead of Click-Based Browsing
Linxi Jiang, Rui Xi, Zhijie Liu +3
This position paper argues that building a reliable agentic Web requires shifting from low-level interaction primitives to typed actions supported by a semantic layer. Today's web…
Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
Tianxin Wei, Noveen Sachdeva, Benjamin Coleman +12
Statefulness is essential for large language model (LLM) agents to perform long-term planning and problem-solving. This makes memory a critical component, yet its management and ev…
Can We Count on LLMs? The Fixed-Effect Fallacy and Claims of GPT-4 Capabilities
Thomas Ball, Shuo Chen, Cormac Herley
In this paper we explore evaluation of LLM capabilities. We present measurements of GPT-4 performance on several deterministic tasks; each task involves a basic calculation and tak…