Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
ORAgentBench: Can LLM Agents Solve Challenging Operations Research Tasks End to End?
Jiajun Li, Mingshu Cai, Yixuan Li +5
Large language models are increasingly deployed as autonomous agents for multi-step tasks in executable environments, yet their ability to perform realistic operations research (OR…
cs.AI2026
MIND-Skill: Quality-Guaranteed Skill Generation via Multi-Agent Induction and Deduction
Yixuan Li, Mingshu Cai, Ziyang Xiao +3
Large language model (LLM) powered AI agents have emerged as a promising paradigm for autonomous problem-solving, yet they continue to struggle with complex, multi-step real-world…