Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
ORAgentBench: Can LLM Agents Solve Challenging Operations Research Tasks End to End?
Jiajun Li, Mingshu Cai, Yixuan Li +5
Large language models are increasingly deployed as autonomous agents for multi-step tasks in executable environments, yet their ability to perform realistic operations research (OR…
cs.AI2026
AutoSci: A Memory-Centric Agentic System for the Full Scientific Research Lifecycle
Weitong Qian, Beicheng Xu, Zhongao Xie +16
Scientific research has traditionally been human-intensive, requiring researchers to coordinate literature, ideas, experiments, manuscripts, and review responses across long projec…
cs.AI2026
LABSHIELD: A Multimodal Benchmark for Safety-Critical Reasoning and Planning in Scientific Laboratories
Qianpu Sun, Xiaowei Chi, Yuhan Rui +5
Artificial intelligence is increasingly catalyzing scientific automation, with multimodal large language model (MLLM) agents evolving from lab assistants into self-driving lab oper…