Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
SOP-Bench: Complex Industrial SOPs for Evaluating LLM Agents
Subhrangshu Nandi, Arghya Datta, Rohith Nama +21
LLM-based agents struggle to execute complex, multi-step Standard Operating Procedures (SOPs) that are fundamental to industrial automation. Existing benchmarks fail to capture the…
cs.AI2025
Design and Evaluation of Cost-Aware PoQ for Decentralized LLM Inference
Arther Tian, Alex Ding, Frank Chen +3
Decentralized large language model (LLM) inference promises transparent and censorship resistant access to advanced AI, yet existing verification approaches struggle to scale to mo…