Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
SOP-Bench: Complex Industrial SOPs for Evaluating LLM Agents
Subhrangshu Nandi, Arghya Datta, Rohith Nama +21
LLM-based agents struggle to execute complex, multi-step Standard Operating Procedures (SOPs) that are fundamental to industrial automation. Existing benchmarks fail to capture the…
cs.AI2024
GrounDial: Human-norm Grounded Safe Dialog Response Generation
Siwon Kim, Shuyang Dai, Mohammad Kachuee +3
Current conversational AI systems based on large language models (LLMs) are known to generate unsafe responses, agreeing to offensive user input or including toxic content. Previou…