3 papers
cs.AI2026
SOP-Bench: Complex Industrial SOPs for Evaluating LLM Agents
Subhrangshu Nandi, Arghya Datta, Rohith Nama +21
LLM-based agents struggle to execute complex, multi-step Standard Operating Procedures (SOPs) that are fundamental to industrial automation. Existing benchmarks fail to capture the…
cs.AI2025
Mind the Goal: Data-Efficient Goal-Oriented Evaluation of Conversational Agents and Chatbots using Teacher Models
Deepak Babu Piskala, Sharlene Chen, Udita Patel +2
Evaluating the quality of multi-turn chatbot interactions remains challenging, as most existing methods assess interactions at the turn level without addressing whether a user's ov…
cs.CL2025
THELMA: Task Based Holistic Evaluation of Large Language Model Applications-RAG Question Answering
Udita Patel, Rutu Mulkar, Jay Roberts +6
We propose THELMA (Task Based Holistic Evaluation of Large Language Model Applications), a reference free framework for RAG (Retrieval Augmented generation) based question answerin…