2 papers
cs.AI2026
SOP-Bench: Complex Industrial SOPs for Evaluating LLM Agents
Subhrangshu Nandi, Arghya Datta, Rohith Nama +21
LLM-based agents struggle to execute complex, multi-step Standard Operating Procedures (SOPs) that are fundamental to industrial automation. Existing benchmarks fail to capture the…
cs.SE2025
Cybernaut: Towards Reliable Web Automation
Ankur Tomar, Hengyue Liang, Indranil Bhattacharya +2
The emergence of AI-driven web automation through Large Language Models (LLMs) offers unprecedented opportunities for optimizing digital workflows. However, deploying such systems…