Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
AgentCE-Bench: Agent Configurable Evaluation with Scalable Horizons and Controllable Difficulty under Lightweight Environments
Wang Yang, Chaoda Song, Xinpeng Li +7
Existing Agent benchmarks suffer from two critical limitations: high environment interaction overhead (up to 41\% of total evaluation time) and imbalanced task horizon and difficul…
cs.AI2026
Empirical-MCTS: Continuous Agent Evolution via Dual-Experience Monte Carlo Tree Search
Hao Lu, Haoyuan Huang, Yulin Zhou +2
Inference-time scaling strategies, particularly Monte Carlo Tree Search (MCTS), have significantly enhanced the reasoning capabilities of Large Language Models (LLMs). However, cur…