7 papers
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation
Inder Preet, Shuxin Lin, Dhaval Patel
Probing the capabilities of Large Language Models (LLMs) and building robust solutions for Multiple-Choice Question Answering (MCQA) remain central challenges in natural language u…
AssetOpsBench: Benchmarking AI Agents for Task Automation in Industrial Asset Operations and Maintenance
Dhaval Patel, Shuxin Lin, James Rayfield +7
AI for Industrial Asset Lifecycle Management aims to automate complex operational workflows, such as condition monitoring and maintenance scheduling, to minimize system downtime. W…
From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents
Ling Yue, Kushal Raj Bhandari, Ching-Yun Ko +6
Large language model (LLM)-based systems are becoming increasingly popular for solving tasks by constructing executable workflows that interleave LLM calls, information retrieval,…
Efficient Embedding-based Synthetic Data Generation for Complex Reasoning Tasks
Srideepika Jayaraman, Achille Fokoue, Dhaval Patel +1
Synthetic Data Generation (SDG), leveraging Large Language Models (LLMs), has recently been recognized and broadly adopted as an effective approach to improve the performance of sm…
SPIRAL: Symbolic LLM Planning via Grounded and Reflective Search
Yifan Zhang, Giridhar Ganapavarapu, Srideepika Jayaraman +3
Large Language Models (LLMs) often falter at complex planning tasks that require exploration and self-correction, as their linear reasoning process struggles to recover from early…
Toward a Trustworthy Optimization Modeling Agent via Verifiable Synthetic Data Generation
Vinicius Lima, Dzung T. Phan, Jayant Kalagnanam +2
We present a framework for training trustworthy large language model (LLM) agents for optimization modeling via a verifiable synthetic data generation pipeline. Focusing on linear…