citation prediction 1idea generation 1large language models 1reinforcement learning 1scientific discovery 1
From the 1 of 12 linked papers with an AI index.
Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
ContextWeave: A Real-World Workflow Benchmark
Bo Wang, Yuqian Yao, Enxi Wang +25
Memory is essential as language agents move from isolated tasks to long-horizon, stateful workflows, yet existing evaluations often reduce it to retrieval or question answering. We…
cs.AI2026
AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers
Tianyu Huai, Tingshuo Fan, Xinchi Chen +5
As LLMs evolve from code completion systems into autonomous scientific agents, evaluating their ability to conduct experiments has become increasingly important. Existing benchmark…
cs.AI2026
AstroReason-Bench: Evaluating Unified Agentic Planning across Heterogeneous Space Planning Problems
Weiyi Wang, Xinchi Chen, Jingjing Gong +2
Recent advances in agentic Large Language Models (LLMs) have positioned them as generalist planners capable of reasoning and acting across diverse tasks. However, existing agent be…