4 papers
VAMPS: Visual-Assisted Mathematical Problem Solving Benchmark
Amirhossein Dabiriaghdam, Shayan Vassef, Mohammadreza Bakhtiari +5
Multimodal large language models are increasingly capable of complex reasoning, yet their performance often degrades when they must externalize a problem through a tool and then re…
SOP-Bench: Complex Industrial SOPs for Evaluating LLM Agents
Subhrangshu Nandi, Arghya Datta, Rohith Nama +21
LLM-based agents struggle to execute complex, multi-step Standard Operating Procedures (SOPs) that are fundamental to industrial automation. Existing benchmarks fail to capture the…
SMARTAPS: Tool-augmented LLMs for Operations Management
Timothy Tin Long Yu, Mahdi Mostajabdaveh, Jabo Serge Byusa +5
Large language models (LLMs) present intriguing opportunities to enhance user interaction with traditional algorithms and tools in real-world applications. An advanced planning sys…
Evaluating LLM Reasoning in the Operations Research Domain with ORQA
Mahdi Mostajabdaveh, Timothy T. Yu, Samarendra Chandan Bindu Dash +5
In this paper, we introduce and apply Operations Research Question Answering (ORQA), a new benchmark designed to assess the generalization capabilities of Large Language Models (LL…