6 papers
ORAgentBench: Can LLM Agents Solve Challenging Operations Research Tasks End to End?
Jiajun Li, Mingshu Cai, Yixuan Li +5
Large language models are increasingly deployed as autonomous agents for multi-step tasks in executable environments, yet their ability to perform realistic operations research (OR…
StarOR: Synergizing Tree Search and Test-Time Reinforcement Learning for Optimization Modeling
Jiajun Li, Yu Ding, Shisi Guan +2
Optimization modeling is inherently hierarchical, requiring a precise sequence of symbolic commitments. Traditional learning-based automated optimization modeling methods improve m…
AutoSci: A Memory-Centric Agentic System for the Full Scientific Research Lifecycle
Weitong Qian, Beicheng Xu, Zhongao Xie +16
Scientific research has traditionally been human-intensive, requiring researchers to coordinate literature, ideas, experiments, manuscripts, and review responses across long projec…
Can LLMs Think Like Consumers? Benchmarking Crowd-Level Reaction Reconstruction with ConsumerSimBench
Tianyu Wang, Jiajun Li, Jianghao Lin
LLMs are increasingly used as ``digital consumers'' to simulate public opinion, pre-test marketing decisions, and anticipate audience response. However, existing evaluations rarely…
AlphaEval: Evaluating Agents in Production
Pengrui Lu, Bingyu Xu, Wenjun Zhang +24
The rapid deployment of AI agents in commercial settings has outpaced the development of evaluation methodologies that reflect production realities. Existing benchmarks measure age…
LABSHIELD: A Multimodal Benchmark for Safety-Critical Reasoning and Planning in Scientific Laboratories
Qianpu Sun, Xiaowei Chi, Yuhan Rui +5
Artificial intelligence is increasingly catalyzing scientific automation, with multimodal large language model (MLLM) agents evolving from lab assistants into self-driving lab oper…