4 papers
ORAgentBench: Can LLM Agents Solve Challenging Operations Research Tasks End to End?
Jiajun Li, Mingshu Cai, Yixuan Li +5
Large language models are increasingly deployed as autonomous agents for multi-step tasks in executable environments, yet their ability to perform realistic operations research (OR…
StarOR: Synergizing Tree Search and Test-Time Reinforcement Learning for Optimization Modeling
Jiajun Li, Yu Ding, Shisi Guan +2
Optimization modeling is inherently hierarchical, requiring a precise sequence of symbolic commitments. Traditional learning-based automated optimization modeling methods improve m…
AlphaEval: Evaluating Agents in Production
Pengrui Lu, Bingyu Xu, Wenjun Zhang +24
The rapid deployment of AI agents in commercial settings has outpaced the development of evaluation methodologies that reflect production realities. Existing benchmarks measure age…
LABSHIELD: A Multimodal Benchmark for Safety-Critical Reasoning and Planning in Scientific Laboratories
Qianpu Sun, Xiaowei Chi, Yuhan Rui +5
Artificial intelligence is increasingly catalyzing scientific automation, with multimodal large language model (MLLM) agents evolving from lab assistants into self-driving lab oper…